ABOUT THE ROLE:You will ensure engineering teams get the right data they need-high-quality tasks and evaluations-by crafting and executing high-value Human Data projects. You sit at the intersection of model teams and data operations: partnering with engineering, designing projects that capture meaningful signals, and driving measurement of data impact. This is a hands-on role for people who understand training and evaluation deeply and want to influence data strategy.
RESPONSIBILITIES:- Partner with model and engineering teams to understand needs and translate them into high-value data projects and evaluation strategies.
- Own end-to-end delivery of critical data and evaluation projects that capture meaningful training signals and support rapid model development.
- Leverage AI agents and existing platforms to measure data effectiveness and quantify impact (data yield, eval lift, usage).
- Maintain rigorous data integrity and truthfulness, including validation processes for factual accuracy; prioritize quality over quantity.
- Research and apply techniques for data collection, annotation, generation, and multi-modal integration.
- Shape Grok's behavior and domain performance through targeted data work; improve annotation workflows using agents and no/low-code approaches where helpful.
- Manage plans that shape model behavior via data management, optimization, and analysis, including resources and timelines.
- Act as a liaison between engineering, technical staff, and tutoring / Human Data teams to drive alignment and knowledge sharing.
- Collaborate with Human Data Ops and stakeholders to scale projects, share learnings, and contribute to demand forecasting; report status, insights, and blockers for rapid decisions.
BASIC QUALIFICATIONS:- Bachelor's degree, or 4+ years of relevant experience in lieu of a degree
- Experience collaborating with cross-functional teams (engineering, research, product, or annotation/operations groups)
- Demonstrated experience analyzing datasets to identify trends, anomalies, quality issues, or integrity problems
PREFERRED SKILLS AND EXPERIENCE:- Degree in engineering, computer science, data analysis, or a related STEM discipline (Bachelor's or higher)
- Direct experience curating, evaluating, or improving training or evaluation datasets for large language models or other AI/ML systems
- Experience designing, supporting, or optimizing annotation workflows or data processes that prioritize factual accuracy and data integrity
- Familiarity with multimodal data (text + images, code, or other modalities) or domain-specific data (science, mathematics, programming, recent events, etc.)
- Experience with model or dataset evaluations focused on quality, truthfulness, or alignment with product goals
- Comfort using AI agents, no/low-code tools, or light scripting (e.g., Python) to prototype workflows and measurement
- Experience with SQL or other data analysis tools
ADDITIONAL REQUIREMENTS:- Weekend work may be required.
- Travel to other SpaceXAI sites may be required.
COMPENSATION AND BENEFITS:$100,000 - $186,000 USD
Base salary is just one part of our total rewards package at SpaceXAI, which also includes equity, comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short & long-term disability insurance, life insurance, and various other discounts and perks.