Position SummaryWe are seeking an AI Research Scientist to lead the development of evaluation frameworks that measure and improve the reasoning capabilities of Large Language Models (LLMs) for investment decision-making. This individual will partner closely with investment professionals, AI researchers, and AI engineers to capture expert reasoning, design rigorous benchmarks, build high-quality evaluation datasets, and develop methodologies that guide model training, post-training, and deployment.
This role sits at the intersection of AI research, investment domain expertise, and applied machine learning, with a primary focus on ensuring our models reason accurately, consistently, and reliably across complex private market investment workflows.
This individual should:
- Think like an AI researcher but build practical systems.
- Rigorously evaluate model reasoning instead of relying solely on benchmark scores.
- Enjoy creating datasets, experiments, and evaluation methodologies.
- Work effectively with domain experts to extract tacit knowledge.
- Independently drive ambiguous research problems from concept through implementation.
ResponsibilitiesLLM Evaluation & Benchmark Development- Design and implement comprehensive evaluation frameworks for investment-focused LLMs.
- Develop quantitative and qualitative metrics to measure reasoning quality, factuality, consistency, calibration, and investment decision quality.
- Build representative benchmark datasets covering real-world investment workflows.
- Design test suites that measure model performance across varying investment scenarios and edge cases.
- Run and analyze evaluation results to identify opportunities to improve model performance.
Data Collection & Dataset Development- Develop strategies for collecting, curating, annotating, and maintaining high-quality training and evaluation datasets.
- Translate investment reasoning into structured datasets suitable for training and evaluation.
- Define annotation guidelines and quality assurance processes.
- Partner with investment professionals to capture expert reasoning and decision-making processes.
Post-Training & Model Improvement- Support supervised fine-tuning (SFT), preference optimization (DPO/RLHF), reinforcement learning, and other post-training methodologies.
- Design evaluation loops that measure gains from post-training efforts.
- Identify model failure modes and recommend improvements.
- Collaborate with AI engineers to deploy evaluation pipelines into production workflows.
Cross-Functional Collaboration- Work closely with investment teams to understand investment processes and reasoning.
- Partner with AI engineers to integrate evaluation systems into model development pipelines.
- Collaborate with AI researchers on new evaluation methodologies and experimental designs.
- Communicate research findings and recommendations to technical and non-technical stakeholders.
Minimum Qualifications- M.S. or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, Statistics, or related field (or equivalent experience).
- Experience developing LLM evaluation frameworks, benchmarks, or experimentation methodologies.
- Experience collecting, curating, annotating, and managing datasets for AI systems.
- Strong understanding of model evaluation metrics, experimental design, and statistical analysis.
- Experience with SFT, DPO/RLHF, reinforcement learning, or related post-training techniques.
- Strong Python programming skills and experience with modern ML frameworks.
- Experience working with foundation models and modern LLM tooling.
- Excellent communication and cross-functional collaboration skills.
Preferred Qualifications- Experience applying AI to finance or investment research.
- Experience building human preference datasets.
- Research in NLP, AI alignment, LLM evaluation, or reasoning.
- Publications in leading AI conferences or journals.
- Ability to rapidly learn complex technical and business domains.
Benefits/CompensationThe compensation range for this role is specific to New York and takes into account a wide range of factors including but not limited to the skill sets required/preferred; prior experience and training; licenses and/or certifications.
The anticipated base salary range for this role is $180,000 to $200,000.
In addition to the base salary, the hired professional will enjoy a comprehensive benefits package spanning retirement benefits, health insurance, life insurance and disability, paid time off, paid holidays, family planning benefits and various wellness programs. Additionally, the hired professional may also be eligible to participate in an annual discretionary incentive program, the award of which will be dependent on various factors, including, without limitation, individual and organizational performance.
Due to the high volume of candidates, please be advised that only candidates selected to interview will be contacted by The Carlyle Group.