Research Engineer, Clinical ReasoningLocation: New York City or San Francisco
Work Style: Hybrid Onsite
Employment Type: Full-time
Focus: Clinical AI, Post-Training, Reinforcement Learning, Evals, Retrieval, Agentic Reasoning
About the RoleOur client is hiring a
Research Engineer, Clinical Reasoning to help build the reasoning systems, learning methods, retrieval algorithms, post-training data pipelines, and evaluation infrastructure behind its clinical AI platform.
This role blends research and engineering. The right person can run experiments, post-train models, build high-quality data pipelines, debug ML systems, and ship production-quality infrastructure that helps clinical AI improve with every iteration.
This is not a pure research role and it is not a generic software engineering role. It is designed for someone with a strong spike in ML or AI research and the engineering ability to turn open-ended research ideas into working systems.
What You'll BuildAgentic Clinical Reasoning
- Design and implement next-generation clinical AI reasoning systems
- Build architectures for reasoning, reflection, verification, tool use, routing, uncertainty handling, and escalation
- Create systems where specialized agents and models work together to support safe, reliable clinical decisions
- Help clinical AI systems earn greater autonomy over time through better reasoning, evaluation, and feedback loops
Evaluation and Measurement
- Build evaluation platforms, rubrics, simulations, and experiments that measure clinical AI performance
- Identify when a benchmark score rewards the wrong behavior
- Determine whether improvements should come from reasoning, retrieval, model behavior, data, or engineering changes
- Measure whether each intervention genuinely improves safety, accuracy, usefulness, and clinical reliability
Models, Learning, and Post-Training
- Run post-training experiments to improve clinical AI behavior
- Apply methods such as fine-tuning, distillation, reinforcement learning, preference optimization, and prompt or system optimization
- Build training data, feedback, reward, and experimentation pipelines
- Partner with in-house researchers to translate training objectives into concrete data and evaluation specifications
- Scale training infrastructure and experimentation workflows
- Create and manage both real-world and synthetic data pipelines
Search, Retrieval, and Grounding
- Build search, ranking, retrieval, and grounding algorithms that connect clinical reasoning to trusted medical evidence, patient context, and partner-specific content
- Improve retrieval based on downstream clinical decision quality, not just document relevance
- Build systems that help clinical AI produce grounded, evidence-backed, trustworthy answers
What We're Looking For- Strong spike in ML or AI research, especially in areas like post-training, reinforcement learning, evaluations, interpretability, model behavior, or agentic systems
- Strong software engineering ability with the ability to build reliable tools, debug systems, and ship working infrastructure
- Experience post-training models and building or curating high-quality post-training data
- Ability to design, evaluate, and improve data pipelines for model training and clinical reasoning
- Strong judgment around data quality, task realism, evaluation design, and model behavior
- Fast problem-solving and code comprehension skills, especially in unfamiliar codebases or ML pipelines
- Comfort debugging PyTorch workflows, model training issues, data pipelines, and evaluation systems
- Ability to translate research goals into concrete datasets, experiments, and measurement frameworks
- Ability to work across research and engineering, turning open-ended AI problems into usable systems
- Clear communication with research, engineering, clinical, product, and data teams
Ideal BackgroundSuccessful candidates may come from backgrounds such as:
- Applied ML research
- Research engineering
- Post-training or reinforcement learning
- Model evaluation and interpretability
- Data science with strong engineering depth
- Software engineering with a research-oriented AI focus
- Clinical AI, healthcare AI, or safety-critical AI systems
The team cares more about depth of work, research judgment, engineering ability, and ability to ship than a specific credential or career path.
Bonus Experience- Published research, patents, meaningful open-source contributions, or novel production ML systems
- Experience building AI systems at an early-stage or high-growth company
- Experience in healthcare, clinical AI, or another regulated or safety-critical domain
- Familiarity with clinical workflows, healthcare data, HIPAA, FHIR, EHR systems, or HL7
- Experience with human-feedback systems, RLHF, simulation, or synthetic data generation
- Experience in AI safety, bias detection, calibration, fairness, or model reliability
- Experience building evaluation harnesses for LLMs, agents, or clinical decision systems
Why This Opportunity- Join a frontier AI lab focused on clinical reasoning
- Work with real patient interactions and real clinical workflows
- Build systems that help clinical AI reason, learn, retrieve evidence, and improve over time
- Own problems end-to-end across research, data, training, evaluation, and deployment
- Partner with researchers, engineers, clinicians, and product teams on high-impact healthcare AI systems
- Build in a safety-critical domain where evaluation, grounding, and reliability truly matter
- Help define how clinical AI earns trust and autonomy over time
Ideal Candidate ProfileThe ideal candidate is a research-minded engineer with a real spike in ML, AI research, post-training, reinforcement learning, evaluations, or model behavior.
They can understand research objectives, build the data and evaluation systems behind them, debug ML pipelines, post-train models, and ship production-quality infrastructure. They have strong data taste, strong engineering ability, and the judgment to build clinical AI systems that are measurable, grounded, and safe.
This person wants to work at the intersection of frontier AI research and real-world healthcare.