Principal Machine Learning Operations Engineer

Definitive Healthcare, US

$158K — $294K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in a relevant field (Computer Science, Data Science, etc.)
  • 12+ years of experience in data and ML engineering
  • 5+ years of experience in MLOps platforms
  • Hands-on experience with Databricks and AWS
  • Expert proficiency in Python, experience with TensorFlow and PyTorch.

Responsibilities

  • Design and maintain scalable ML and data pipelines for CI/CD and ML deployments.
  • Architect and manage enterprise feature store and vector store for low-latency feature retrieval.
  • Build and manage complex workflow orchestration for model training and data processing.
  • Serve as subject matter expert for Generative AI infrastructure and frameworks.
  • Establish efficient environments for model training and fine-tuning open-source models.
  • Optimize model inference pipelines for high throughput and low latency.
  • Implement real-time monitoring systems for model performance and system health.

Benefits

  • Comprehensive medical, dental, and vision coverage
  • Unlimited paid time off
  • 401(k) plan with employer contribution
  • Participation in company bonus or commission plan.
Full Job Description
About the Role

We are seeking a hands-on Principal MLOps Engineer & Functional Lead to architect, build, and scale our next-generation Machine Learning and Generative AI infrastructure. In this role, you will design a unified, automated, and secure AI ecosystem capable of training and deploying both traditional predictive models and advanced Generative/Agentic AI systems at enterprise scale.

As the Functional Lead, you will drive engineering excellence, establish MLOps best practices across the organization, mentor senior engineers, and translate strategic business objectives into robust, highly scalable automated pipelines.

What You'll Do
  • Design, implement, and maintain scalable, robust ML pipelines and data pipelines for CI/CD, and ML model deployments.
  • Architect and govern the enterprise feature store and vector store infrastructure to support low-latency feature retrieval for both traditional ML and Retrieval Augmented Generation (RAG) applications.
  • Build and manage enterprise-grade workflow orchestration layers to schedule, monitor, and manage complex, multi-stage model training and data processing dependencies.
  • Be a subject matter expert for Generative AI infrastructure, building scalable frameworks to support advanced LLM applications, RAG pipelines, and multi-agent Agentic AI workflows.
  • Establish highly efficient environments for model training and parameter-efficient fine-tuning of open-source and proprietary models.
  • Optimize high-throughput, low-latency model inference pipelines, utilizing advanced caching and compute distribution techniques to handle large-scale concurrent requests.
  • Establish and enforce lifecycle management policies utilizing an enterprise model registry to manage version control, lineage, and tracking from experimentation to production.
  • Architect high-performing, cost-efficient infrastructure utilizing autoscaling groups on AWS to dynamically handle varying training and inference workloads.
  • Implement comprehensive, real-time endpoint monitoring systems to track model performance, data drift, latency, and system health metrics.
  • Serve as the functional lead for ML engineering, defining coding standards, architectural blueprints, framework selections, and operational SLAs.
  • Provide technical guidance, code reviews, and mentorship to senior machine learning and data engineers, fostering an agile culture of innovation, automation, and continuous improvement.


What You'll Bring
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Data Science, or a closely related quantitative engineering discipline.
  • 12+ years of professional experience in data engineering, ML engineering with a minimum of 5 years dedicated to building, scaling, and managing MLOps platforms in production.
  • Deep, hands-on production experience architecting with in Databricks and native Aws environments.
  • Expert proficiency in Python, with strong production experience deploying models built in TensorFlow and PyTorch.

Compensation and Benefits

The salary range for this position is $158,000 - $294,000 per year, which represents the base pay the company reasonably and in good faith expects to pay for this role. Actual pay within this range will be determined based on factors such as relevant experience, skills, and qualifications.

Depending on the position, employees may also be eligible to participate in a company bonus or commission plan. All employees are eligible for a comprehensive benefits package, including medical, dental, and vision coverage, unlimited paid time off, and participation in the company's 401(k) plan with employer contribution.

Similar Jobs

More Jobs at Definitive Healthcare, US

More Information Technology Jobs

Find similar Principal Machine Learning Operations Engineer jobs: