Turing

Staff Research Scientist, STEM

Turing • $250K — $400K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • PhD or equivalent in a technical field such as machine learning, computer science, or engineering.
  • Proven ability to formulate and execute original research.
  • Deep understanding of modern LLMs and frontier AI trends.
  • Strong skills in experimental design and quantitative reasoning.
  • Adept at digesting complex technical literature and acquiring new knowledge rapidly.
  • Proficient in Python, capable of building research prototypes and evaluation systems.
  • Exceptional technical writing and communication skills.

Responsibilities

  • Identify gaps in benchmark and evaluation literature.
  • Design innovative benchmarks across STEM fields.
  • Develop evaluations for emerging model capabilities.
  • Create methodologies for task generation, grading, and validation.
  • Generate high-quality synthetic training data for STEM applications.
  • Examine factors influencing performance from synthetic data.
  • Research reliability and verification issues in AI systems.
  • Evaluate workflows for AI agent collaboration in scientific research.

Benefits

  • Work in a dynamic, research-first environment.
  • Collaborate with leading experts in frontier AI research.
  • Leverage significant autonomy to propose new research directions.
  • Contribute to impactful research outputs that enhance the AI field.
Full Job Description
The Role

Turing is seeking exceptional Staff Research Scientists to join our STEM research organization and develop new ways to evaluate, train, and improve frontier AI systems.

This is a research-first role focused on problems where the right benchmark, dataset, or methodology often does not yet exist. You will identify important gaps in the literature, propose ambitious new research directions, and take projects from initial hypothesis through experimentation, benchmark construction, and publication.

Our research is deliberately focused on frontier STEM evaluation, synthetic data, hallucination and reliability, and agentic science. We are looking for scientists who can recognize important problems early, formulate them precisely, and design rigorous research programs to answer them.
What You'll Do
Frontier benchmarks and evaluation
  • Identify high-impact gaps in existing benchmark and evaluation literature.
  • Design novel benchmarks in and across STEM fields and on general model functionality.
  • Develop evaluations for emerging model capabilities that are poorly captured by traditional static benchmarks.
  • Design rigorous task-generation, grading, contamination-control, difficulty-calibration, and validation methodologies.
  • Build benchmarks that can become both valuable research contributions and meaningful standards for evaluating frontier models.
Synthetic data and post-training
  • Develop methods for generating high-quality synthetic STEM training data.
  • Study how task selection, difficulty, diversity, verification, filtering, and data quality affect downstream performance.
  • Explore methods for generating useful training signal in domains where expert human data is scarce or expensive.
  • Design experiments that determine when synthetic data genuinely improves capabilities rather than simply increasing training volume.
Hallucination, reliability, and verification
  • Study hallucination, uncertainty, calibration, and epistemic failure in technical domains.
  • Develop evaluations and methods for improving factual reliability, self-correction, verification, citation, and appropriate abstention.
  • Investigate when models should reason internally, invoke tools, seek external evidence, or recognize that they do not know.
Agentic science
  • Research AI systems capable of performing extended scientific and technical work.
  • Develop workflows involving literature search, coding, simulation, tool use, experimentation, verification, and iterative reasoning.
  • Evaluate long-horizon scientific agents and identify the bottlenecks preventing them from reliably performing real research.
  • Explore new approaches to human-AI and multi-agent scientific collaboration.
New research directions

The areas above are our core focus, not an exhaustive list. Researchers will also have significant latitude to propose new programs in areas such as reasoning, model evaluation, AI-for-science, data generation, and emerging capabilities.
What We're Looking For
  • PhD or equivalent research experience in machine learning, computer science, mathematics, physics, chemistry, biology, engineering, statistics, or another highly technical field.
  • Demonstrated ability to formulate and execute original research.
  • Strong understanding of modern LLMs and the frontier AI research landscape.
  • Excellent experimental design, quantitative reasoning, and scientific judgment.
  • Ability to rapidly understand unfamiliar technical literature and develop expertise in new areas.
  • Strong Python skills and the ability to independently build research prototypes and evaluation pipelines.
  • Excellent technical writing and communication.
  • Comfort working in a fast-moving environment where the research agenda evolves with the frontier.

A strong publication record is valuable, but we care most about whether you can identify important questions, design rigorous ways to answer them, and execute quickly enough for the results to matter.
What Success Looks Like

You might:
  • Identify a major capability that existing benchmarks fail to measure and create the benchmark that becomes the standard for evaluating it.
  • Discover a failure mode in current synthetic-data pipelines and develop a method that materially improves post-training.
  • Build a new evaluation that changes how frontier labs understand hallucination, reasoning, or scientific capability.
  • Develop an agentic workflow that substantially advances performance on complex scientific research tasks.
  • Launch an entirely new research direction that grows into a major program within Turing.
This role is required to be in office five days a week, based in any of Turing's offices in San Francisco, Palo Alto, or Seattle.

Compensation: $250,000 to $400,000 OTE + Equity

About Turing

Turing is a technology company that provides a platform for companies to hire remote software developers. The company's platform uses artificial intelligence to match companies with developers who have the skills and experience they need. Turing was founded in 2018 and is headquartered in San Francisco, California. The company has raised $32 million in funding to date.
Learn more about Turing
Size
200 employees
Industry
Founded
2018

Similar Jobs

More Jobs at Turing

More Information Technology Jobs

Find similar Staff Research Scientist, STEM jobs: