Snorkel AI

Senior/Staff FDE - Synthetic Data Generation

Snorkel AI$180K — $320K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in machine learning engineering, data science, or applied AI
  • Proficient in Python and reliable production data systems
  • Hands-on experience with LLMs and the GenAI stack
  • Strong understanding of ML evaluation metrics
  • Experience in synthetic data and data augmentation
  • Expertise in LLM evaluation techniques
  • Demonstrated problem-solving from definition to delivery

Responsibilities

  • Design scalable synthetic data generation and evaluation pipelines
  • Translate model objectives into synthetic data strategies
  • Develop ML-assisted workflows for training datasets
  • Build automated evaluators and quality measurement frameworks
  • Measure synthetic data impact on model performance
  • Lead technical workstreams from design through delivery
  • Communicate technical tradeoffs to stakeholders

Benefits

  • Employee stock options
  • Health, dental, and vision insurance
  • Flexible work arrangements
  • Professional development opportunities
  • Access to cutting-edge technology and tools
Full Job Description
About the Role

Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives.

In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream model outcomes.

You will work across the full delivery lifecycle-from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements.
Main Responsibilities
Synthetic Data Generation & Evaluation
  • Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases
  • Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications
  • Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases
  • Build automated evaluators, quality checks, and measurement frameworks to assess correctness, relevance, diversity, coverage, and adherence to customer requirements
  • Design and run experiments to measure the impact of synthetic data on downstream model performance and iteratively improve generation approaches
  • Package and deliver production-grade datasets with standardized formats, quality assurance, and clear documentation
Forward Deployed Engineering & Customer Partnership
  • Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions
  • Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value
  • Rapidly prototype and productionize solutions across models, data pipelines, APIs, and custom applications
  • Communicate technical tradeoffs, experimental results, and recommendations clearly to technical and cross-functional stakeholders
  • Serve as a trusted technical partner to customers and internal delivery teams, resolving complex blockers and driving alignment
Technical Leadership & Scale
  • Identify recurring patterns across customer engagements and turn successful solutions into reusable pipelines, evaluators, tooling, and best practices
  • Define and improve technical standards for synthetic data generation, experimentation, evaluation, and delivery
  • Partner with DaaS Engineering and Product teams to influence platform and product capabilities based on real-world customer needs
  • Lead technical design reviews, share expertise, and provide guidance to other engineers
  • Stay current with emerging synthetic data, LLM evaluation, and data curation techniques and assess their applicability to customer problems
What We're Looking For
  • 5+ years of experience in machine learning engineering, data science, applied AI, forward deployed engineering, or a similar technical role
  • Strong Python skills and experience building reliable production data or ML systems, including containerizing with Docker and deploying on cloud platforms (e.g., AWS, GCP, or Azure)
  • Hands-on experience with LLMs-building model-based applications and data workflows with the modern GenAI/LLM stack, and integrating systems, models, and data sources through APIs
  • Strong understanding of ML experimentation and evaluation, including defining metrics and using empirical results to guide technical decisions
  • Experience building synthetic data, data augmentation, or model-generated training and evaluation datasets
  • Experience with LLM evaluation techniques, including LLM-as-a-judge, model-based evaluation, rubric-based evaluation, or custom evaluators
  • Demonstrated ability to take ambiguous technical problems from problem definition through delivery, with strong technical communication and experience working directly with customers and cross-functional stakeholders
  • Experience serving as a technical lead-setting technical direction, driving architecture and key decisions, mentoring engineers, and creating reusable approaches that influence broader engineering or product outcomes
Preferred Qualifications
  • Experience developing datasets for fine-tuning, preference optimization, or benchmarking, including human-in-the-loop generation and review workflows
  • Experience building agentic environments and tasks-repo-scale coding tasks, tool-agent-user interaction design, and agent tool protocols and interoperability
  • Experience with reinforcement learning for LLMs, including reward and verifier design and RL with verifiable rewards (RLVR)
  • Experience working in fast-paced, customer-facing environments where requirements and technical approaches evolve quickly
Compensation

The base salary range for this position is $180,000-$320,000, with an additional variable compensation opportunity. The exact mix of base salary and variable compensation will depend on the role level and work location. Final compensation will be determined based on job-related skills, experience, relevant education or training, interview performance, and other business considerations.

All offers also include equity in the form of employee stock options, as well as benefits.

Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.

Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.

Salary range(s) for this role

$180,000-$320,000 USD

About Snorkel AI

Snorkel AI is an artificial intelligence company that provides a platform for building and managing machine learning models. The company was founded in 2019 and is headquartered in San Francisco, California. Snorkel AI's platform is designed to make it easier for developers and data scientists to create and manage machine learning models, using a technique called programmatic labeling. The company's platform is used by a number of large enterprises, including Intel, Google, and Microsoft. Snorkel AI has raised over $50 million in funding to date.
Learn more about Snorkel AI
Size
50 employees
Industry
Founded
2019

Similar Jobs

More Jobs at Snorkel AI

More Information Technology Jobs

Find similar Senior/Staff FDE - Synthetic Data Generation jobs: