The AI Engineering team is responsible for working closely with the research and modeling teams to create state-of-the-art NLP models for specific tasks, and deploy them in a production setting designed to serve our customers at scale. We are looking for a Machine Learning Engineer to help build and evaluate the core intelligence behind our agentic AI systems. This role will play a key part in designing and owning evaluation frameworks that ensure quality, safety, and performance across complex agentic systems.
We're looking for a Lead Machine Learning Engineer to own and grow the evaluation platform that measures quality, safety, and performance across ASAPP's agentic AI systems- the infrastructure that tells us, with confidence, whether a model or agent change is actually an improvement before it reaches customers.
This a hybrid role with 10-12 days of in-office presence per month to balance flexibility with collaboration.
What you'll do- Help develop the technical roadmap and architecture for the evaluation platform, from offline benchmarking to online/production monitoring of agentic and LLM-based systems.
- Design eval methodologies appropriate to different stages of the pipeline: golden/regression test sets, human-in-the-loop review workflows, LLM-as-judge approaches, and automated metrics for task success, safety, and hallucinations.
- Build the data infrastructure evaluation depends on: annotation and labeling pipelines, dataset versioning, data quality checks, and tooling that lets researchers and product teams run and interpret experiments without needing platform team help.
- Partner closely with Research, Product, and Platform teams to productize experiments into robust AI solutions
- Represent the eval platform to stakeholders outside the immediate team- set expectations on what "good" looks like for a model/agent release, and report on platform health and coverage.
- Stay current with advancements in ML, NLP, voice, and LLM systems, and contribute actively to technical discussions across teams.
- Mentor and support other engineers through design reviews, feedback, and knowledge sharing.
What you'll need- Deep, hands-on experience building and operating evaluation systems for modern ML/LLM/agentic systems- not just consuming existing eval tools.
- Demonstrated experience leading the technical direction of a project or small team: setting architecture, driving design reviews, and being accountable for a system's long-term health (not just shipping features).
- Strong architectural skills, with proven experience designing complex, data-intensive software systems and production experience with Python, AWS, Kubernetes, and/or Docker.
- Experience designing data pipelines for ML evaluation- labeling/annotation workflows, dataset versioning and quality control, and reproducible benchmarking.
- A Bachelor's Degree in CS or other related fields
- Demonstrated technical mentorship of junior and mid-level engineers, driving adoption of best practices and architectural alignment for scalability and extensibility.
- Desire to learn, teach, and collaborate closely with cross-functional peers.
What we'd like to see- Experience building and evaluating agentic systems at scale.
- Experience with voice/audio quality evaluations.
- Production experience with LLM-centric services (e.g., inference, orchestration, evaluation, monitoring)
- Familiarity with large-scale ML experimentation, benchmarking, or simulation frameworks.
- Experience with conversational/customer-support AI domains (e.g., containment rate, conversation quality, goal completion).
- Knowledge of techniques for optimizing model architectures for faster inference.
- Experience with AWS, CI/CD, Kafka, Athena
$170,000 - $190,000 a year
Compensation package also includes a performance bonus on top of the listed salary range
Separately, we also offer a compelling equity grant comprised of stock options
Benefits include:
Competitive compensation with stock options
Comprehensive medical, vision, and dental insurance
401k matching
Fitness and wellness stipend
Mental well-being benefits
Professional learning and development stipend
Parental leave, including adoptive and foster parents
3 weeks paid time off (increases with tenure) along with sick leave, bereavement and jury duty