Member of Technical Staff - Audio and Voice AI

Stuut, Inc.

$130K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in software engineering, 2+ years in applied AI/ML, speech, or audio systems
  • Experience in deploying production-ready voice or conversational AI systems
  • Proficient in speech-to-text, text-to-speech, and audio processing technologies
  • Integrated and fine-tuned LLMs for conversational systems
  • Familiarity with LLMOps / MLOps best practices
  • Fluent in Python with experience in PyTorch, TensorFlow, or audio ML frameworks
  • Experience in building real-time systems and understanding performance tradeoffs
  • Ability to translate business needs into scalable AI solutions

Responsibilities

  • Build and deploy production-grade voice AI systems
  • Craft high-quality, responsive voice user experiences
  • Adapt and fine-tune audio and multimodal models for optimal performance
  • Engineer scalable AI pipelines for audio ingestion and processing
  • Establish evaluation frameworks to measure AI system performance
  • Automate financial workflows through AI voice technologies
  • Collaborate with cross-functional teams to develop effective solutions
  • Measure impact and refine systems based on customer feedback

Benefits

  • Medical, dental & vision insurance coverage
  • 401(k) plan with matching
  • Equity options
  • Flexible PTO policy
  • Parental leave
Full Job Description
The RoleWe're hiring a Member of Technical Staff - Audio and Voice AI Systems to design, build, and deploy AI-powered voice and audio systems that solve real-world financial operations challenges. You'll take state-of-the-art research in speech, audio, and multimodal AI and translate it into production-grade, real-time voice experiences that deliver measurable customer impact.

From intelligent voice agents that interact with customers to audio-driven workflow automation and speech-based data extraction, you'll create scalable, reliable AI systems that integrate seamlessly into Stuut's platform. Your work will directly shape how finance teams and their customers interact with Stuut through voice.

This is a hands-on role for an engineer who thrives at the intersection of audio AI, real-time systems, and practical business impact-turning cutting-edge models into trusted, delightful voice experiences for enterprise finance workflows.

What You'll Do
  • Build & Deploy Voice AI Systems: design and ship production-ready audio and voice-based AI features, including real-time voice agents and speech-driven workflows.
  • Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial use cases.
  • Adapt & Fine-Tune Audio and Multimodal Models: fine-tune and optimize speech, audio, and LLM-based models for accuracy, latency, and reliability in real-world environments.
  • Engineer Real-Time, Scalable AI Pipelines: build end-to-end AI/ML pipelines spanning audio ingestion, streaming inference, orchestration, and monitoring with enterprise-grade availability and performance.
  • Establish Evaluation & Monitoring Frameworks (LLMOps): design rigorous evaluation systems to measure quality, latency, accuracy, drift, and business outcomes for voice and text-based AI systems.
  • Automate Financial Workflows via Voice: develop AI-powered voice automations that reduce manual effort in collections, reconciliation, and customer communication.
  • Collaborate Cross-Functionally: partner with Product, Engineering, Design, and customers to translate business needs into effective, user-centered voice AI solutions.
  • Measure & Communicate Impact: define success metrics and continuously improve AI systems based on real-world usage and customer feedback.


You Might Be a Fit If You...
  • Have 5+ years of software engineering experience, with 2+ years focused on applied AI/ML, speech, or audio systems in production.
  • Have built and shipped voice, audio, or conversational AI systems used by real customers.
  • Have experience with speech-to-text, text-to-speech, audio processing, or multimodal models.
  • Have integrated and fine-tuned LLMs for conversational or agent-based systems.
  • Understand LLMOps / MLOps best practices, including deployment pipelines, monitoring, evaluation, and A/B testing.
  • Are fluent in Python and experienced with PyTorch, TensorFlow, Transformers, or audio ML frameworks.
  • Have built real-time or low-latency systems and understand the tradeoffs involved.
  • Can translate business and UX requirements into robust, scalable AI solutions.
  • Have experience integrating AI systems into existing enterprise or SaaS platforms.
  • Enjoy working on ambiguous problems where product definition, UX, and engineering meet.

Compensation
  • Top-of-market salary and equity package
  • Benefits (for U.S.-based full-time employees)
  • Medical, dental & vision insurance coverage for you
  • 401(k) & Match
  • Equity
  • Flexible PTO
  • Parental Leave

Similar Jobs

More Jobs at Stuut, Inc.

More Enterprise Technology Jobs

Find similar Member of Technical Staff - Audio and Voice AI jobs: