Job Title: AI Researcher Lead 1
City: San Francisco
State/Province: California
Posting Start Date: 9/21/26
Job Description:
Job Description
Role Overview As an AI Data Foundry, we partner with top-tier labs to build the critical SFT and RLHF datasets that train the next generation of LLMs, SLMs, and Physical AI. We are seeking a Lead Researcher who focuses on the rigorous science of measurement and alignment. In this role, you will architect preference data pipelines, design proactive benchmarks from the ground up, and publish novel methodologies on arXiv and at top-tier conferences (e.g., NeurIPS, ICLR).
͏
Key Responsibilities
• Design Alignment Methodologies: Architect frameworks for human-in-the-loop and synthetic data generation. Design pairwise preference rubrics (utilizing Bradley-Terry models) and reward mechanisms for DPO, KTO, and PPO pipelines.
• Build Proactive Benchmarks: Lead the creation of novel, domain-specific benchmarks, spanning agentic reasoning to physical AI, that evaluate complex capabilities rather than relying on saturated, static datasets.
• Decontaminate & Validate: Build robust pipelines to detect data contamination (via n-gram overlap, embedding similarity) to ensure deliverables are mathematically sound and mitigate reward hacking.
• Research Publication: Conduct independent research on evaluation frameworks and alignment science, co-authoring papers to establish our technical authority in the field.
• Technical Leadership: Work closely with domain-specific researchers to synthesize ground-truth data into cohesive, scoreable, and statistically sound evaluation frameworks.
͏
Candidate Profile
• Academic Background: Advanced degree (PhD or high-impact MSc) in Computer Science, Artificial Intelligence, Mathematics, or a highly quantitative field.
• Technical Expertise: Deep familiarity with modern alignment techniques (RLHF, DPO, RLAIF) and standard evaluation harnesses (e.g., EleutherAI LM Eval, HELM, OpenAI Evals).
• Core Stack: Advanced proficiency in PyTorch, HuggingFace, vLLM, and complex data manipulation.
• Model Experience: Hands-on experience fine-tuning or evaluating open-weight models (Llama 3, Mistral, Qwen) with a deep understanding of their architectural bottlenecks.
͏
͏
Expected annual pay for this role ranges from $240,000.00 to $375,000.00. Based on the position, the role is also eligible for Wipro's standard benefits including a full range of medical and dental benefits options, disability insurance, paid time off (inclusive of sick leave), other paid and unpaid leave options