About the RoleThis role sits at the intersection of technical program management, data operations, and vendor coordination for an early-stage AI/ML infrastructure company. You will own complex data and evaluation programs for frontier AI labs and internal teams, turning ambiguous technical asks into well-structured, executable programs. Your work directly improves the quality of LLM and agent evaluation at scale.
What You'll Do- Own data and evaluation programs end-to-end, from initial scoping and requirements gathering through production, quality assurance, delivery, and retrospective.
- Translate ambiguous requests from AI labs and internal teams into clear specifications with defined milestones, owners, dependencies, and acceptance criteria.
- Maintain and improve data quality procedures using quantitative signals and qualitative inspection, including analysis of task coverage, difficulty, and reward distributions.
- Manage external vendors throughout the delivery lifecycle and improve delivery pipelines at scale.
- Streamline data production logistics by improving workflows, documentation, automation, and cross-functional handoffs.
- Partner with research, platform, and marketplace teams to translate learnings and feedback into product improvements.
What We're Looking For- 3+ years owning complex technical programs in AI/ML, data operations, support engineering, applied research, or similarly cross-functional environments.
- Demonstrated ability to translate ambiguous technical requirements into executable specifications with clear milestones, owners, dependencies, and acceptance criteria.
- Strong quantitative and qualitative judgment about data, with the ability to move between aggregate metrics and individual examples to assess dataset quality.
- Experience maintaining and improving complex data quality procedures using both quantitative metrics and qualitative inspection methods.
- Experience managing external vendors and scaling delivery pipelines.
- Background supporting frontier AI labs, technical enterprise customers, or research teams with urgent and evolving requirements.
- Experience with RLHF, reinforcement-learning environments, agent evaluations, human-in-the-loop data pipelines, annotation workflows, or expert data collection.
- Strong written and verbal communication skills, with the ability to turn complex technical information into clear decisions and next steps.
- Early-stage startup experience and comfort working independently in fast-paced environments.
Compensation & BenefitsCompetitive compensation package including equity. Full benefits include 100% employer-covered medical, dental, and vision insurance, a 401k, commuter benefits, and unlimited access to leading AI productivity tools. Visa sponsorship and relocation support are available for strong full-time candidates.
LocationOn-site in San Francisco, CA or Singapore. Remote candidates who can maintain 70 to 80% time zone overlap with either office location are also welcome to apply.