Lead Research Engineer, Data Quality

Clera

$150K — $250K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in research or data quality engineering, focused on AI/ML data evaluation.
  • Proven experience leading teams or projects in data quality or AI/ML evaluation.
  • Advanced skills in Python programming, Docker, and Linux environments.
  • Deep understanding of high-quality training data attributes: realistic, learnable, diverse, reliable.
  • Experience turning research insights into functioning production systems and data pipelines.
  • Expertise in validating synthetic data methods at scale.

Responsibilities

  • Lead the data quality team to create systems evaluating RL environments, synthetic data, and benchmarks.
  • Define and implement the company's data quality strategy, including QC systems and standards.
  • Develop innovative validation methods for synthetic data, including failure-mode analysis and auditing.
  • Collaborate with research engineers and data vendors to enhance data generation workflows.
  • Translate qualitative insights into practical tools, dashboards, and validation systems.
  • Cultivate an internal research culture that promotes high-quality training data attributes.
  • Mentor team members to uphold standards of technical excellence and rapid execution.

Benefits

  • Early-stage equity participation.
  • Visa sponsorship available.
Full Job Description
About the Role

We're an early-stage AI/ML infrastructure company (11-50 people) building the platform that AI labs and businesses use to create, manage, and scale reinforcement learning (RL) environments and high-quality post-training datasets. We're looking for a Lead Research Engineer, Data Quality to own the strategy and systems that measure, improve, and scale training data for frontier agents.

In this role, you'll lead the data quality team, shape our internal research culture, and define what makes agent training data truly useful - not just superficially correct. You'll work on-site in San Francisco, CA. Visa sponsorship is available.
What You'll Do
  • Lead the data quality team in building systems that evaluate thousands of tasks across RL environments, synthetic data, benchmarks, and domain-specific workflows.
  • Define the company's data quality strategy - build QC systems, enforce standards, and design experiments to grade agent outputs.
  • Develop novel methods for validating synthetic data at scale, including failure-mode analysis, task mutation checks, and trajectory auditing.
  • Partner with research engineers, domain experts, and data vendors to diagnose quality issues and improve data generation workflows.
  • Translate qualitative research insights into production systems: internal tools, dashboards, validation pipelines, and feedback loops.
  • Help build internal research taste around what makes agent training data realistic, learnable, diverse, reliable, and useful.
  • Mentor research engineers to maintain a high bar for technical rigor, clarity, and execution speed.
What We're Looking For

Required (dealbreakers):
  • 5+ years of experience in a research or data quality engineering role, specifically building systems for AI/ML data evaluation.
  • Demonstrated experience leading technical projects or teams in data quality or AI/ML evaluation.
  • Advanced proficiency in Python, Docker, and Linux environments.

Also required:
  • Ability to reason deeply about characteristics of high-quality training data (realistic, learnable, diverse, reliable) for AI agents.
  • Experience translating research insights into production systems and data pipelines (e.g., validation pipelines, feedback loops).
  • Experience developing and implementing methods for validating synthetic data at scale.
  • Experience collaborating with research engineers, domain experts, and data vendors to diagnose quality issues.
  • Track record mentoring engineers on technical rigor and execution speed.
  • Comfort navigating complex systems involving domain experts, vendors, generated data, model outputs, graders, and infrastructure.

Nice to have:
  • Experience leading teams on ambiguous technical projects from problem definition through implementation and iteration.
  • Experience working with subject-matter experts to convert domain judgment into scalable review or generation systems.
  • Experience designing metrics, experiments, and QA/QC processes from scratch.
  • Prior early-stage startup experience and comfort operating independently in fast-paced environments.
  • Strong written communication skills - able to explain methodology clearly to researchers, engineers, and external audiences.
Compensation & Benefits
  • Salary: $150,000 - $250,000 USD annually, depending on experience.
  • Early-stage equity participation.
  • Visa sponsorship available.
Location

This is an on-site role based in San Francisco, CA. Candidates must be willing to work in-office. Visa sponsorship is available for qualified candidates.

Similar Jobs

More Jobs at Clera

  • Founding Forward Deployed Engineer
    $110K — $135K *
    Austin, TX 78745 (Travis County)
    Enterprise Technology
    In-Person
  • GTM Engineer
    $120K — $200K *
    Los Angeles, CA 90011 (Los Angeles County)
    Enterprise Technology
    In-Person
  • Founding Engineer
    $120K — $250K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Founding GTM
    $120K — $200K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Forward Deployed Engineer
    $150K — $250K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person

More Information Technology Jobs

Find similar Lead Research Engineer, Data Quality jobs: