Office Hours

Software Engineer- Benchmarking

Office Hours$160K — $210K *
US-AnywhereRemote in United States
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of professional software engineering experience with a focus on Python.
  • Strong capability in preparing and maintaining datasets for accurate comparisons.
  • Experience with containerization using Docker to build reproducible environments.
  • Ability to collaborate with researchers to translate methodologies into functional systems.
  • Familiarity with AI evaluation frameworks is a plus.

Responsibilities

  • Prepare and maintain benchmark datasets through cleaning and validation.
  • Build and maintain evaluation pipelines for consistent model assessments.
  • Create containerized environments for evaluating model performance.
  • Support fine-tuning of small open-source LLMs and analyze performance.
  • Develop scoreboards and leaderboards for published evaluation results.
  • Create analysis tools to identify model failure modes and performance tracking.
  • Collaborate with teams to ensure data accuracy and integration in publications.

Benefits

  • Medical, dental, and vision coverage.
  • Monthly wellness and fitness stipend.
  • Paid time off along with company holidays.
  • Annual company off-sites in various locations.
  • Parent-friendly policies, including remote flexibility and paid family leave.
Full Job Description
Software Engineer, Benchmarking (SF, NYC, or Remote)

About the role

We're looking for a Software Engineer to build and run the platform behind our AI model evaluations. You'll work closely with our research team to prepare benchmark datasets, build the pipelines and environments our evaluations run in, and turn results into published output.

Our researchers design the methodology. You'll turn it into systems that run consistently and reproducibly, so results stay comparable across models, agent scaffolds, and time. Evaluations draw on the knowledge domains our expert network covers.

What you'll do
  • Prepare and maintain benchmark datasets: Own the data work behind our benchmarks, including cleaning, preparation, conversion into runnable formats, and ongoing maintenance. Validate that tasks are complete, consistent, and executable, and flag ambiguities that would compromise results.
  • Build and maintain evaluation pipelines: Build the infrastructure that runs evaluations consistently across model APIs and terminal agents, so results are reproducible and comparable.
  • Build evaluation environments: Create lightweight, containerized environments and viewers for tasking and for evaluating model performance on tool use.
  • Support model experiments: Help fine-tune small open-source LLMs and compare baseline against post-training performance.
  • Develop the scoreboard and leaderboard: Build the published views of our results, including model-level, benchmark-level, task-level, domain-level, and rubric-level performance.
  • Build analysis tools: Make it easy to identify recurring failure modes, compare models and agent scaffolds fairly, and track capability improvements and regressions over time.
  • Collaborate: Work closely with researchers and engineers to make sure evaluation data and outputs are accurate, consistent, and well integrated into what we publish.
  • Build tooling for data creation and review: Support expert annotation and data-generation projects by building lightweight HTML viewers and internal tools for task authoring, review, quality control, and structured data collection.
What you bring
  • Solid engineering skills: 4+ years of professional experience building and maintaining complex systems, with strong Python. You write robust, maintainable code and are comfortable diving deep into existing codebases and infrastructure.
  • Data rigor: Experience preparing, cleaning, and maintaining datasets, and the care to make sure two results are genuinely comparable.
  • Comfort with containers and environments: Experience with Docker and building reproducible execution environments.
  • Collaborative: You work well alongside researchers and scientists and can translate their methodology into working systems.

Hands-on experience running AI evaluations, or with frameworks like Harbor, Terminal-Bench, or Inspect, is a strong plus.

Tech Stack
  • Evaluations: Python, model APIs, agent/evaluation frameworks, custom evaluation tooling
  • Models: APIs from the major AI providers, terminal agents, and open-source models via the Hugging Face ecosystem and PyTorch
  • Environments: Docker
  • Publishing: React, Next.js, Tailwind
  • Workflow: GitHub, Slack, Notion, Linear
Bonus Experience
  • Experience fine-tuning or post-training open-source LLMs, or other hands-on machine learning work
  • Experience with agentic, multi-turn, long-context, or tool-use evaluation
  • Experience validating LLM-as-judge or rubric-based grading setups
  • Background or strong interest in a scientific or technical domain
  • Experience building data-heavy dashboards, leaderboards, or visualizations
  • Open-source contributions or published work related to benchmarks and measurement
Benefits + Perks
  • Competitive salary and equity
  • Medical, dental, and vision coverage
  • 401(k)
  • Monthly wellness and fitness stipend
  • Paid time off policy, along with company holidays
  • Annual company off-sites (Tahoe, Mendocino, Mexico City, San Diego, Park City)
  • Parent-friendly policies, remote flexibility, and paid family leave

Pay Transparency Notice

Full-time offers include base salary, equity, and benefits.

Pay range: $160,000-$210,000, based on seniority, relevant experience and location

This role can be fully remote or hybrid out of our SF or NYC offices.

Don't meet every single requirement? Studies have shown that some candidates, especially underrepresented groups such as women and people of color, are less likely to apply to jobs unless they meet every single qualification. At Office Hours we believe in building a diverse and inclusive workplace, so if you're excited about this role but don't meet every qualification in the job description, we still encourage you to apply. You could still be the right candidate for this or other roles at Office Hours!

About Office Hours

Office Hours is a software company that provides a scheduling and appointment booking platform for businesses. The platform allows businesses to manage their appointments and schedules, and provides customers with an easy way to book appointments online. Office Hours' solutions are designed to improve efficiency and reduce administrative overhead for businesses. The company was founded in 2019 and is headquartered in Santa Monica, California.
Learn more about Office Hours
Size
10 employees
Industry
Founded
2019

Similar Jobs

More Jobs at Office Hours

More Information Technology Jobs

Find similar Software Engineer- Benchmarking jobs: