Snorkel AI

Senior | Staff Software Engineer - AI / ML

Snorkel AI • $208K — $315K *
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years building production ML or software systems from prototype to production
  • Hands-on experience with LLM or ML workloads in production
  • Deep grounding in statistics and experimentation, including hypothesis testing and confidence intervals
  • Strong Python skills with solid software engineering fundamentals
  • Experience designing evaluations and interpreting results rigorously
  • Proactive problem identification and clear communication skills with diverse teams

Responsibilities

  • Develop efficient methodologies for evaluating long-horizon agents
  • Implement AI model routing for cost-effective evaluations
  • Fine-tune and serve open-weight models aligned with frontier quality
  • Build models predicting task difficulty for frontier systems
  • Create golden datasets and maintain quality measurement for AI data
  • Transform research prototypes into reusable components for production use

Benefits

  • Contribute to solving frontier problems in AI and ML
  • Focus exclusively on open ML or LLM challenges without underlying infrastructure issues
  • Significant influence in establishing ML engineering practices at Snorkel
  • Visible impact on the speed and quality of frontier AI data generation
Full Job Description
The role

Frontier AI data is expensive to make and hard to measure. Every task we deliver is tested against the strongest models, often through many long-running agent rollouts. Your job is to make that process faster, cheaper, and more rigorous with ML and AI

You will be one of the early members of ML & Research Engineering at Snorkel. You will study how frontier-grade data is generated and evaluated, form hypotheses, validate them against real production data, and ship the winners at scale. You will shape the discipline's direction, its standards, and the team that grows around it.
What you'll work on
  • Efficient agentic evals. Cut the cost of long-horizon agent evaluation with adaptive sampling, statistically grounded early stopping, model cascades, caching, and cheap-first gating.
  • AI model routing. Route every eval and judge call to the cheapest model that clears the quality bar, with fallback, monitoring, and cost attribution.
  • Fine-tuned small models. Fine-tune and serve open-weight models (LoRA and other parameter-efficient methods) where they match frontier quality, and know when they don't.
  • Predictive difficulty. Build models that estimate how hard a task is for frontier systems before running a single rollout.
  • Measurement for AI data. Build golden datasets, quantify the accuracy and calibration of LLM-as-judge systems, and make quality reproducible across projects.
  • Research to production. Turn research prototypes into reusable, configurable components that forward deployed engineers and researchers use on every project.
What you'll bring
  • 5+ years building production ML or software systems, with end-to-end ownership from prototype to production
  • Hands-on experience running LLM or ML workloads in production, and comfort reasoning about non-deterministic systems
  • Deep grounding in statistics and experimentation: experiment design, hypothesis testing, sampling, and confidence intervals
  • Strong Python and software engineering fundamentals, including testing, code review, and system design
  • Experience designing evaluations and interpreting results rigorously
  • A habit of finding high-impact problems before they are assigned, and clear communication with researchers, engineers, and business partners
Nice to have
  • Fine-tuning and serving open-weight models, and judging when a smaller model meets the quality bar
  • Building LLM evaluation or experimentation platforms, model gateways, or routing systems
  • Experience with agentic workloads, benchmarks, or RL environments
  • A record of taking research into production: publications, open-source work, or shipped research-driven features
  • MS or PhD in Computer Science, Machine Learning, Statistics, or a related field
Why join now
  • Frontier problems. Measure and shape the tasks designed to challenge the strongest models in the world.
  • All AI, no plumbing. Every problem on this team is an open ML or LLM problem.
  • Founding impact. Help define what ML engineering means at Snorkel and grow the team that carries it forward.
  • Visible results. Your work shows up directly in the speed, quality, and cost of the data frontier AI is built on.


Actual compensation will be determined based on factors including skills, qualifications, experience, and geographic location.

Salary range(s) for this role

$208,000-$315,000 USD

About Snorkel AI

Snorkel AI is an artificial intelligence company that provides a platform for building and managing machine learning models. The company was founded in 2019 and is headquartered in San Francisco, California. Snorkel AI's platform is designed to make it easier for developers and data scientists to create and manage machine learning models, using a technique called programmatic labeling. The company's platform is used by a number of large enterprises, including Intel, Google, and Microsoft. Snorkel AI has raised over $50 million in funding to date.
Learn more about Snorkel AI
Size
50 employees
Industry
Founded
2019

Similar Jobs

More Jobs at Snorkel AI

More Consumer Technology Jobs

Find similar Senior | Staff Software Engineer - AI / ML jobs: