Staff Applied Scientist, Reinforcement Learning

Hippocratic AI

$160K — $190K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • MS or PhD in Computer Science or relevant field
  • 5+ years of experience in Natural Language Processing (NLP), large language model (LLM) training, or reinforcement learning (RL)
  • 2+ years of experience specifically in RL for LLM post-training
  • Experience with large-scale LLM training (50 billion+ parameters, multi-node)
  • Strong programming skills in Python and PyTorch
  • Familiarity with RLHF, RLVR, and LLM-as-judge methodologies

Responsibilities

  • Design and develop post-training methods for reinforcement learning and on-policy distillation
  • Build and assess reward models and verification systems
  • Create conversational AI environments and simulations for healthcare training
  • Automate post-training processes with intelligent agents
  • Conduct experiments to analyze factors improving post-training metrics
  • Work collaboratively with research, engineering, and clinical teams

Benefits

  • Opportunity to work in an innovative healthcare environment
  • Collaborative team culture with in-office engagement
  • Involvement in meaningful projects impacting patient care
  • Access to advanced technologies and methodologies in AI
  • Potential for professional growth and continuing education opportunities
Full Job Description
Location Requirement

We believe the best ideas happen together. To support fast collaboration and a strong team culture, this role is expected to be in our Palo Alto office five days a week, unless otherwise specified.

About the Role

LLM post-training is where raw capability becomes reliable, safe behavior - and in healthcare, the stakes are as high as they get. You'll own the Reinforcement Learning (RL) and On-Policy Distillation (OPD) post-training pipeline end to end, to improve our models' clinical reasoning, safety, and alignment. Your models will be deployed to interact with millions of patients across diverse clinical use cases.
What You'll Do
  • Design RL and OPD post-training methods (RLHF, RLVR, OPD, etc.)
  • Build and evaluate reward models, verifiers, and LLM-as-judge pipelines
  • Develop conversational AI environments and simulations for healthcare RL training with synthetic data
  • Automate post-training loops with agents (auto-research)
  • Run rigorous experiments to understand what drives post-training gains
  • Collaborate with research, engineering, and clinical teams
What You Bring
  • MS or PhD in CS or relevant field
  • 5+ years or experience in NLP, LLM training, or RL
  • 2+ years experience in RL for LLM post-training
  • Experience with large-scale (50B+ parameter and multi-node) LLM training
  • Strong Python and PyTorch coding skills
  • Experience with RLHF, RLVR, LLM-as-judge or similar methods for LLM post-training


Nice-to-Have:
  • Publications at top venues (NeurIPS, ICML, ICLR, ACL, EMNLP)
  • Healthcare domain experience

Please be aware of recruitment scams impersonating Hippocratic AI. All recruiting communication will come from [redacted].com email addresses. We will never request payment or sensitive personal information during the hiring process.

Similar Jobs

More Jobs at Hippocratic AI

More Information Technology Jobs

Find similar Staff Applied Scientist, Reinforcement Learning jobs: