OpenAI

Machine Learning Engineer, Core Experimentation

OpenAI • $130K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in machine learning and software engineering
  • Proven track record in developing end-to-end ML products
  • Strong skills in Python for building production systems
  • Experience with various ML processes: dataset design, evaluation, deployment
  • Depth in areas like LLMs, ranking systems, or causal ML
  • Solid understanding of experimentation and statistical reasoning
  • Ability to translate partner needs into actionable roadmaps

Responsibilities

  • Set and execute the technical roadmap for ML capabilities
  • Build systems that synthesize and learn from historical experiments
  • Develop predictive models and simulation workflows to assess impact
  • Create datasets and retrieval pipelines with high data quality
  • Establish rigorous evaluation methods for model performance
  • Turn insights into actionable product and workflow solutions
  • Collaborate with cross-functional teams on causal inference and decision-making

Benefits

  • Collaborative in-person work environment in Bellevue
  • Opportunity to shape the future of ML in product decision-making
  • Focus on impact-driven ML products and real-world decision improvements
  • Engagement in building sophisticated internal tools and capabilities
  • Joining a growing team dedicated to high-quality production ML
Full Job Description
About the Role

We are looking for a Machine Learning Engineer to lead the technical direction for ML-powered experimentation and insights capabilities. You will build production systems that learn from privacy-protected product and experimentation data to generate evidence-backed insights and support decision-making, and help teams decide which ideas are worth testing live.

This is an end-to-end, 0-to-1 role. You will work across ML modeling, retrieval and LLM systems, statistical methods, simulation, data and training pipelines, backend services, and user- and agent-facing product experiences. The hard part is not merely producing a plausible answer. It is making each insight and prediction traceable, calibrated, useful, and safe enough to influence real product decisions.

Live experiments remain the source of causal validation. You will design systems that make uncertainty explicit, backtest against historical outcomes, compare predictions with online results, learn from misses, and abstain when the evidence is weak. You will preserve clear review, permission, and approval boundaries as automation becomes more powerful.

You will collaborate closely with teams building ChatGPT, Codex, model measurement workflows, consumer products, Growth, business subscription experiences, developer products, and shared infrastructure. You will turn their most important learning and decision problems into general platform capabilities that can support the full company.

In This Role, You Will
  • Set and execute the technical roadmap for Generative Insights and Predictive Experimentation, from early prototypes through production adoption.
  • Build cross-experiment learning systems that retrieve and synthesize historical experiments, detect recurring effects and segment behavior, reanalyze prior results when data or methods improve, and generate hypotheses with clear evidence and provenance.
  • Develop predictive models and simulation workflows, including simulation-based evaluation approaches, to estimate likely impact, affected segments, regression risk, and uncertainty before a full live experiment.
  • Create high-quality datasets and feature or retrieval pipelines from exposures, events, metrics, experiment metadata, and replay data, with strong lineage, freshness, privacy, and data-quality controls.
  • Establish rigorous evaluation through offline benchmarks, backtests, calibration, drift monitoring, prediction-to-outcome comparisons, and explicit failure or abstention behavior.
  • Turn models into durable product, API, and agent workflows that move from an insight to experiment design, approval-gated action, and measured learning.
  • Partner deeply with data science and product teams on experiment design, causal inference, sequential decision-making, variance reduction, and the boundary between prediction and causal evidence.
  • Build reliable services and intuitive workflows so sophisticated ML capabilities are understandable and useful to teams making high-stakes product decisions.
  • Provide technical leadership across engineering, product, data science, and research partners, and raise the bar for production ML quality across the platform.


You Might Thrive In This Role If You
  • Have led ambiguous 0-to-1 production ML products where success was measured by better real-world decisions, not only offline model metrics.
  • Have strong hands-on experience across the ML lifecycle: dataset design, training or adaptation, evaluation, deployment, monitoring, and iteration.
  • Bring depth in one or more of LLM and retrieval systems, ranking or recommendation, forecasting or anomaly detection, causal ML or experiment analysis, or simulation. You do not need to have done all of them.
  • Have strong software engineering fundamentals and can build high-quality production systems in Python while working comfortably across data, backend, and platform boundaries.
  • Have a strong grounding in machine learning, statistics, computer science, or a related field through formal study or equivalent practical experience.
  • Understand experimentation and statistical reasoning, especially why predictive accuracy is not the same as causal validity.
  • Treat calibration, uncertainty, provenance, privacy, and human review as product requirements, not cleanup work.
  • Can translate ambiguous partner questions into a product and technical roadmap, and work well with product, data science, research, and infrastructure partners.
  • Enjoy building for internal power users and agents, and can make sophisticated ML capabilities feel clear and actionable.
  • Value in-person collaboration and want to help shape a growing Bellevue-based team.


Location and Workplace

This role is based in Bellevue, Washington. The team works in person and uses that time to move quickly, solve ambiguous problems together, and stay close to the product teams we support.

About OpenAI

OpenAI is an artificial intelligence research laboratory consisting of the for-profit corporation OpenAI LP and its parent company, the non-profit OpenAI Inc. The company was founded in 2015 by a group of technology leaders, including Elon Musk, Sam Altman, Greg Brockman, Ilya Sutskever, and John Schulman. OpenAI's mission is to develop and promote friendly AI for the betterment of humanity. The company has developed a number of cutting-edge AI technologies, including GPT-3, a language processing system that can generate human-like text. OpenAI has received funding from a number of high-profile investors, including LinkedIn co-founder Reid Hoffman and venture capitalist Peter Thiel.
Learn more about OpenAI
Size
100 employees
Industry
Founded
2015

Similar Jobs

More Jobs at OpenAI

More Information Technology Jobs

Find similar Machine Learning Engineer, Core Experimentation jobs: