Research Engineer, Synthetic Data

Clera

$150K — $250K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-4 years in software engineering, ML engineering, or AI research with relevant project experience.
  • Hands-on expertise in building end-to-end synthetic data pipelines for AI/ML.
  • Proficient in Python and comfortable in Linux with tools like Docker.
  • Strong grasp of synthetic data quality criteria, evaluation metrics, and their limitations.
  • Experience implementing evaluation frameworks for AI models or large language models.
  • Skilled in creating automated systems for dataset generation and processing at scale.
  • Ability to independently manage and deliver projects with minimal guidance.

Responsibilities

  • Build and maintain a comprehensive synthetic data pipeline for AI training tasks.
  • Collaborate closely with experts to develop realistic synthetic tasks across various domains.
  • Design methods for generating diverse and learnable synthetic tasks.
  • Develop tools to mutate, validate, and scale synthetic tasks efficiently.
  • Analyze performance metrics of models on synthetic tasks to identify learnings and improvements.
  • Establish metrics to measure the diversity, realism, learnability, and quality of synthetic tasks.

Benefits

  • Visa sponsorship available.
  • Equity participation in an early-stage startup.
Full Job Description
About the Role

We're a ~15-person engineering team - made up of Olympiad medalists and published researchers - building infrastructure that aligns AI to real-world workflows through reinforcement learning environments and post-training data. We're hiring Research Engineers to own the synthetic data pipeline: transforming domain-specific workflows into scalable, high-quality training tasks for AI agents.

This is a high-ownership, low-bureaucracy role. You'll be working in genuinely unstructured problem spaces where the roadmap is yours to define. Visa sponsorship is available.
What You'll Do
  • Build and maintain the end-to-end synthetic data pipeline, converting domain-specific workflows into realistic, structured, and challenging training tasks for AI agents.
  • Collaborate with subject-matter experts to generate synthetic tasks across professional and technical domains.
  • Design synthetic task generation methods that produce diverse, realistic, and learnable outputs.
  • Build tooling to mutate, validate, and iteratively improve synthetic tasks at scale.
  • Analyze model and agent performance on synthetic tasks to understand what they teach and where they break down.
  • Develop metrics to quantify synthetic task diversity, realism, learnability, and overall quality.
What We're Looking For

Required:
  • 2-4 years of experience in software engineering, ML engineering, or AI research - with a track record of shipping data pipelines, ML infrastructure, or synthetic data systems.
  • Hands-on experience applying synthetic data research methods to build end-to-end data generation pipelines for AI/ML applications.
  • Proficiency in Python; comfortable working in Linux environments with containerization tools such as Docker.
  • Demonstrated understanding of synthetic data quality criteria and evaluation metrics (diversity, realism, learnability) and their limitations - from production or research work.
  • Experience designing, implementing, or maintaining evaluation frameworks, benchmarks, or testing environments for AI agents or large language models.
  • Experience building automated systems to generate, validate, mutate, or process structured datasets at scale.
  • Proven ability to independently own and deliver technical projects end-to-end with minimal predefined requirements.

Nice to Have:
  • Experience detecting edge cases, inconsistencies, or quality issues in synthetic or algorithmically generated datasets.
  • Experience creating synthetic tasks, data, or evaluations across multiple distinct professional or technical domains.
  • Familiarity with reinforcement learning training paradigms, agentic AI workflows, or LLM post-training pipelines.

You'll thrive here if you:
  • Reason from first principles about task design, scoring, and failure modes.
  • Are detail-oriented and naturally spot subtle inconsistencies in data and systems.
  • Are energised by early-stage, ambiguous environments rather than frustrated by them.
  • Communicate clearly and collaborate effectively across time zones.
Compensation & Benefits
  • Salary: $150,000 - $250,000 USD annually
  • Visa sponsorship available
  • Equity participation (early-stage startup)
Location

This role is on-site in San Francisco, CA. Candidates based in or willing to relocate to San Francisco are strongly preferred. The team also has a presence in Singapore.

Similar Jobs

More Jobs at Clera

  • Senior ML/AI Engineer
    $170K — $230K *
    New York, NY 10025 (New York County)
    Enterprise Technology
    In-Person
  • Tech Lead
    $240K — $260K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Founding Agentic Engineer
    $150K — $300K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Founding Full Stack Engineer
    $135K — $155K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person
  • Electrical Engineer
    $100K — $150K *
    Austin, TX 78745 (Travis County)
    Aerospace & Defense
    In-Person

More Information Technology Jobs

Find similar Research Engineer, Synthetic Data jobs: