Two Sigma Investments, LLC

Post-Training Research Scientist

Two Sigma Investments, LLC$165K — $300K *
Finance & Insurance
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS in Science, Technology, Engineering or Math; MS preferred.
  • 1-10 years of experience at a leading AI lab (OpenAI, Anthropic, etc.).
  • Proven track record in deploying RLHF, DPO, or similar systems.
  • Strong grasp of distributed training infrastructure and multi-node GPU setups.
  • Experience managing large compute budgets and experiment designs.
  • Publications in alignment or reward modeling preferred.
  • Proficient in PyTorch/JAX and distributed frameworks.

Responsibilities

  • Lead post-training initiatives for large language models in financial applications.
  • Design and implement RLHF and DPO methods at scale.
  • Develop infrastructure for reward modeling and data collection.
  • Shape research agenda linking post-training methods to quantitative finance.
  • Collaborate with quantitative researchers on task distributions.
  • Resolve issues in production systems reliant on post-training setups.

Benefits

  • Fully paid medical and dental insurance for employees and dependents.
  • Competitive 401k match and employer-paid life & disability insurance.
  • Onsite gyms, wellness activities, and casual dress environment.
  • Tuition reimbursement and sponsorship for conferences and training.
  • Generous vacation policy and unlimited sick days, with caregiver leave options.
  • Flexible hybrid work policy with home office budget.
Full Job Description
Position Summary

We are applying large language models and transformer-based architectures to problems where ground truth is delayed, noisy, and non-stationary. Our systems generate code, run experiments, and iterate autonomously, and we are looking to go beyond supervised fine-tuning.

We are hiring a Post-Training Research Scientist to build RLHF, DPO, and reward modeling capabilities from the ground up. This is a greenfield role: you will define the infrastructure, research agenda, and evaluation frameworks for aligning LLMs to sophisticated, multi-step workflows in a domain where the reward signal is fundamentally different from existing research on human preference or deterministic task completion.

This hire will help own methodology across training, fine-tuning, context management, and model evaluation. You will shape not only the post-training capability but the broader research direction of the team.

You will take on the following responsibilities:

  • Lead post-training efforts for LLMs applied to financial time series and quantitative reasoning
  • Design and execute RLHF, DPO, and related alignment methods at scale, including deployment of substantial compute budgets (O($100mm))
  • Build infrastructure for preference data collection, reward modeling, and policy optimization on financial datasets
  • Drive research agenda connecting post-training methods to quantitative finance applications
  • Collaborate with quant researchers to define task distributions and evaluation frameworks
  • Unblock production systems dependent on post-training capabilities


You should possess the following qualifications:

  • BS or equivalent work experience in Science, Technology, Engineering or Math (an MS is a plus).
  • Minimum 1 year of experience required; 1-10 years of experience preferred (ideally 1-5 years) at a frontier AI lab (OpenAI, Anthropic, DeepMind, Meta FAIR, or equivalent)
  • Shipped post-training systems in production: RLHF, DPO, RLAIF, or related methods
  • Deep understanding of distributed training infrastructure: multi-node GPU clusters, training stability, checkpointing
  • Track record managing large-scale compute: budgeting, experiment design, ablations
  • Publications or demonstrated expertise in alignment, preference learning, or reward modeling
  • Hands-on implementation skills: PyTorch/JAX, distributed frameworks (DeepSpeed, FSDP, etc.)


You will enjoy the following benefits:
  • Core Benefits: Fully paid medical and dental insurance premiums for employees and dependents, competitive 401k match, employer-paid life & disability insurance
  • Perks: Onsite gyms with laundry service, wellness activities, casual dress, snacks, game rooms
  • Learning: Tuition reimbursement, conference and training sponsorship
  • Time Off: Generous vacation and unlimited sick days, competitive paid caregiver leaves
  • Hybrid Work Policy: Flexible in-office days with budget for home office setup


The base pay for this role will be between $165,000 and $300,000. This role may also be eligible for other forms of compensation and benefits, such as a discretionary bonus, health, dental and other wellness plans and 401(k) contributions. Discretionary bonus can be a significant portion of total compensation. Actual compensation for successful candidates will be carefully determined based on a number of factors, including their skills, qualifications and experience.

About Two Sigma Investments, LLC

Two Sigma Investments is a quantitative investment management firm that uses data science and technology to identify investment opportunities. The company's solutions are designed to help investors make better decisions and generate higher returns. Two Sigma Investments offers a range of products, including hedge funds, private equity, and venture capital. The company was founded in 2001 and is headquartered in New York City.
Learn more about Two Sigma Investments, LLC
Size
1,500 employees
Industry
Founded
2001

Similar Jobs

More Jobs at Two Sigma Investments, LLC

More Finance & Insurance Jobs

Find similar Post-Training Research Scientist jobs: