Research Engineer, Synthetic Data

Clera

$150K — $250K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-4 years of experience in software engineering, ML engineering, or AI research roles focused on data pipelines and synthetic data systems.
  • Hands-on experience in building end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python and expertise in Linux with containerization tools like Docker.
  • Solid grounding in synthetic data quality criteria including diversity, realism, and learnability.
  • Experience in creating evaluation frameworks and testing environments for AI agents or large language models.
  • Proven ability to manage technical projects independently and deliver results with minimal guidance.
  • Strong communication skills for effective collaboration in remote settings.

Responsibilities

  • Build end-to-end synthetic data pipelines that generate challenging training tasks for AI agents.
  • Collaborate with experts to develop diverse synthetic tasks across various professional domains.
  • Design task generation methods to create realistic and learnable training examples.
  • Develop tools to validate and improve the quality of synthetic tasks iteratively.
  • Analyze performance outcomes of models and agents on synthetic tasks to identify learning failures.
  • Establish metrics to evaluate task diversity, realism, and overall quality.

Benefits

  • Visa sponsorship available.
  • Opportunity to work within a team of Olympiad medalists and published researchers.
Full Job Description
About the Role

This is a Research Engineer role focused on building synthetic data pipelines for AI agent training, sitting within a ~15-person engineering team of Olympiad medalists and published researchers. You'll design generation methods, validation systems, and quality metrics that directly expand model capabilities - work that sits at the frontier of RL-based AI alignment.
What You'll Do
  • Build end-to-end synthetic data pipelines that transform domain-specific workflows into structured, challenging training tasks for AI agents.
  • Collaborate with subject-matter experts to develop synthetic tasks spanning professional and technical domains.
  • Design task generation methods that produce diverse, realistic, and learnable training examples.
  • Build tooling to mutate, validate, and iteratively improve synthetic task quality.
  • Analyze model and agent performance on synthetic tasks to understand learning outcomes and failure modes.
  • Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
What We're Looking For
  • 2-4 years of experience in software engineering, ML engineering, or AI research roles delivering data pipelines, ML infrastructure, or synthetic data systems.
  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python and experience developing in Linux environments using containerization tools such as Docker.
  • Demonstrated understanding of synthetic data quality criteria and evaluation metrics - diversity, realism, learnability - and their inherent limitations.
  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
  • Track record of independently owning and delivering technical projects end-to-end with minimal predefined requirements.
  • Ability to detect edge cases, inconsistencies, and quality issues in synthetic or algorithmically generated datasets.
  • Comfort operating in unstructured, early-stage environments and reasoning from first principles.
  • Strong communication skills for effective remote collaboration across time zones.
  • Nice to have: Familiarity with reinforcement learning paradigms, agentic AI workflows, or LLM post-training pipelines.
Compensation & Benefits

Salary range: $150,000 - $250,000 USD annually. Visa sponsorship is available.
Location

On-site in San Francisco, CA, United States. Singapore-based candidates are also considered.

Similar Jobs

More Jobs at Clera

More Information Technology Jobs

Find similar Research Engineer, Synthetic Data jobs: