Research Engineer, Benchmarks

Clera

$150K — $250K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-4 years of experience in software engineering, ML engineering, or research roles.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Experience building environments, evaluations, or benchmarks for AI systems.
  • Published research or technical writing on public benchmarks, model failure modes, or evaluation methodology.
  • Deep understanding of realistic and reliable benchmark development.

Responsibilities

  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents.
  • Partner with subject-matter experts to define realistic workflows and tasks for evaluations.
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and analyses to measure benchmark difficulty, reliability, and failure modes.
  • Validate benchmark performance against real-world evaluations and customer needs.

Benefits

  • Visa sponsorship available.
  • Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team.
Full Job Description
About the Role

Join a small, technically elite team - including International Olympiad medalists and published AI researchers - building high-quality benchmarks to evaluate frontier AI agents on realistic, domain-specific workflows. As a Research Engineer, Benchmarks, you'll own the design and implementation of evaluations that frontier labs and enterprise customers trust. This role is central to ensuring our benchmarks are rigorous, credible, and tightly aligned with real-world agent performance.

This is an on-site role based in San Francisco, CA. Visa sponsorship is available.
What You'll Do
  • Design, implement, and own the quality of internal benchmarks for evaluating frontier agents on domain-specific tasks.
  • Partner with subject-matter experts to define realistic workflows and tasks for domain-specific evaluations.
  • Build reliable infrastructure to run models and agents against benchmark tasks at scale.
  • Develop metrics and analyses that measure benchmark difficulty, reliability, and failure modes.
  • Validate that benchmark performance correlates with real-world evaluations, customer needs, and frontier lab expectations.
  • Write clear documentation and benchmark reports that make results legible and credible to technical audiences.
What We're Looking For

Required
  • 2-4 years of experience in software engineering, ML engineering, or research roles.
  • Strong proficiency in Python, Docker, and Linux environments.
  • Experience building environments, evaluations, or benchmarks for AI systems.
  • Published research or technical writing on topics such as public benchmarks, model failure modes, or evaluation methodology.
  • Deep understanding of what makes a benchmark realistic, reliable, and practically useful.
  • Curiosity and genuine ability to understand how real-world workflows operate across diverse domains.
  • Strong attention to detail - a habit of spotting subtle inconsistencies and edge cases in task design.
  • Ability to reason from first principles about task design, scoring, and failure modes.
  • Comfort thriving in unstructured problem spaces and working independently in fast-paced, early-stage environments.
  • Excellent communication skills for collaborating across time zones and with technical teams.
Compensation & Benefits
  • Salary: $150,000 - $250,000 USD annually, depending on experience.
  • Visa sponsorship available.
  • Opportunity for significant early-stage equity and career growth within a high-impact, research-driven team.
Location

This is a full-time, on-site position in San Francisco, CA. Candidates must be willing and able to work in-office.

Similar Jobs

More Jobs at Clera

  • GTM Engineer
    $100K — $150K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Research Engineer, QC Automation
    $150K — $250K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Forward Deployed Research Engineer
    $150K — $250K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Growth Lead
    $100K — $150K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Marketing & Events Lead
    $100K — $150K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person

More Information Technology Jobs

Find similar Research Engineer, Benchmarks jobs: