Aircall

Machine Learning Engineer (Evals and Voice Models)

Aircall$181K — $250K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS in Computer Science, Machine Learning, Statistics, or related field
  • 3+ years of experience in ML Engineering or Applied ML, with 8+ years overall
  • Strong experience evaluating supervised, unsupervised, LLMs, and deep learning models
  • Hands-on experience in failure analysis and evaluating LLMs
  • Experience building automated evaluation systems
  • Strong communication skills for diverse audiences
  • Experience training or fine-tuning voice/speech models (TTS, ASR, speech-to-speech)

Responsibilities

  • Design and document comprehensive evaluation frameworks for voice, chat, and messaging AI agents
  • Train and fine-tune voice models to improve accuracy and naturalness
  • Assess AI-generated solutions through training pipelines and optimizations
  • Analyze system designs to identify strengths and potential failure points
  • Design workflows for human-labeled evaluation data and calibrate evaluation systems against human raters
  • Build and maintain live quality monitoring for deployed AI agents
  • Define metric contracts for published AI metrics and build offline regression suites before deployments

Benefits

  • Opportunity to join Aircall at a pivotal growth stage
  • Commitment to work-life balance
  • Fast-learning environment with an entrepreneurial spirit
  • Diverse team with over 45 nationalities
  • Competitive salary package and benefits
Full Job Description
Aircall's AI suite includes an AI Voice Agent and AI Messaging Agent that autonomously handle calls, WhatsApp, and SMS, plus AI Assist, which delivers real-time coaching, call summaries, and CRM automation for sales and support teams. We are looking for someone that can build out the evaluation foundation across all of these products and other agentic products. You'll work on voice models, agent capability evals, benchmark design, LLM-as-judge systems, failure analysis, and the infrastructure that ties it together by establishing shared metrics, test sets, and tooling to measure accuracy, resolution quality, and safety consistently across products. You will set up repeatable pipelines for regression testing and benchmarking as models and features evolve so teams can ship confidently without re-inventing evaluation methodology for each product. Key Responsibilities • Design and document comprehensive evaluation frameworks for Aircall's AI agents across voice, chat and messaging. • Train and fine-tune voice models (TTS, ASR, speech-to-speech) using production and synthetic data, iterating on architecture, data mix, and training strategy to improve accuracy, naturalness, and latency. • Assess AI generated solutions across training pipelines, experimentation setups, debugging processes, and optimization strategies. • Analyze system design decisions and identify strengths, weaknesses, and potential failure points. • Design annotation guidelines and workflows for human-labeled evaluation data, and calibrate LLM-as-judge systems against human raters to ensure automated evals stay trustworthy over time. • Build and maintain live quality monitoring for deployed AI agents, tracking accuracy, resolution rate, and safety signals in production, and flagging model or data drift before it impacts customers. • Own the metric contract for every published AI metrics, including definition, population, grain, rollup, validity window. • Build release gates, the offline regression suite each AI surface must pass before a prompt, model, or config change ships, measuring reliability across repeated trials, not just average pass rates. • Build voice-specific evaluation: simulated callers across accents, languages, background noise, barge-in, DTMF, and tool failures, with latency and ASR accuracy as first-class quality metrics. Minimum Qualifications • BS in Computer Science, Machine Learning, Statistics, or related field • 3+ years of experience in ML Engineering or Applied ML with 8+ years of overall experience • Strong experience in evaluating supervised, unsupervised, LLMs and deep learning models. • Hands-on experience in failure analysis and evaluating LLMs • Experience building automated evaluation systems • Strong communication skills to articulate complex technical concepts across technical and non-technical audiences • Hands-on experience training or fine-tuning voice/speech models (TTS, ASR, or speech-to-speech), including data pipeline construction and experimentation. Preferred Qualifications • MS / PhD in Computer Science, Machine Learning, Statistics, or related field • Experience evaluating LLMs or agentic systems (e.g., LLM-as-a-judge, RAG evaluation) • Experience with synthetic data generation and prompt engineering • Experience training or fine-tuning voice models at scale, with familiarity in synthetic data generation, model distillation, or low-latency inference optimization for production voice agents. Base salary range: $181,000-$250,000 USD Why join us? Key moment to join Aircall in terms of growth and opportunities Our people matter, work-life balance is important at Aircall Fast-learning environment, entrepreneurial and strong team spirit 45+ Nationalities: cosmopolite & multi-cultural mindset Competitive salary package & benefits

About Aircall

Aircall is a cloud-based phone system and call center software that helps businesses manage their customer support and sales calls. The company's software allows users to make and receive calls from anywhere in the world, using a computer or mobile device. Aircall's software also includes features such as call routing, call recording, and analytics, which help businesses improve their customer service and sales performance. The company was founded in 2014 and is headquartered in New York City.
Learn more about Aircall
Size
500 employees
Industry
Founded
2014

Similar Jobs

More Jobs at Aircall

More Consumer Technology Jobs

Find similar Machine Learning Engineer (Evals and Voice Models) jobs: