MLOps Engineer

SpreeAI Corporation

$145K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience building or operating an ML platform (Ray, Kubeflow, Airflow, or custom solutions)
  • Proficiency in Docker and Kubernetes for orchestration
  • Strong background in Python or Go for infrastructure tooling
  • Hands-on experience with experiment tracking and data versioning at scale
  • Ability to lead platform direction and mentor junior engineers

Responsibilities

  • Design and operate a training-as-a-service infrastructure for ML Scientists
  • Implement CI/CD for models to ensure production quality with automated eval gates
  • Manage experiment tracking and data versioning for large, frequently changing datasets
  • Monitor the health of training jobs and manage cost governance as scaling occurs
  • Collaborate with ML Scientists to improve workflows and platform functionalities
  • Evaluate and integrate external model providers into the ML platform

Benefits

  • Flexible hybrid work model based in San Francisco
  • Opportunity to influence platform design and operations for ML teams
  • Collaboration with a team focused on innovative AI applications like Video Try-On
  • Potential for career growth in a fast-paced, early-stage environment
Full Job Description
About the role

This role owns the ML lifecycle platform end to end: training pipelines, experiment tracking, CI/CD for models, monitoring, and data versioning, so ML Scientists can launch, monitor, and iterate on training runs without managing infrastructure directly. As the Science team expands into Video Try-On and AI Sizing, this platform is what keeps that research moving fast without breaking.

What you'll do

  • Design and operate training-as-a-service infrastructure: a scientist should be able to launch a multi-GPU training job, track metrics, and get notified on completion without touching infra directly
  • Build CI/CD for models: automated eval gates that block a bad checkpoint from reaching production, canary rollout, A/B testing hooks
  • Own experiment tracking (Weights & Biases, MLflow, or Neptune) and data versioning (DVC, LakeFS, or Delta Lake) for datasets in the terabytes that change weekly
  • Monitor training job health (GPU utilization, loss curves, OOM detection) and drive cost governance (spot instances, preemptible VMs, budget alerts) as training costs scale
  • Partner directly with ML Scientists to translate workflow pain points (reproducibility, experiment comparison, checkpoint recovery) into platform abstractions
  • Evaluate and integrate external model providers (SPREEAI uses Byteplus and Fireworks.AI as MaaS providers) into the training and eval platform

What you'll bring

  • Experience building or operating an ML platform: training orchestration (Ray, Kubeflow, Airflow, or a custom solution), Docker, Kubernetes (Jobs/CronJobs, Helm)
  • Python or Go for pipeline orchestration and infrastructure tooling
  • Real experience with experiment tracking and data versioning tools at production scale
  • Comfort owning platform direction, not just executing tickets, and mentoring engineers as the team scales
  • Comfort with broad ownership across the ML lifecycle in an early-stage, fast-moving environment

You'll thrive here if

You treat ML Scientists as your customers and actively design the boundary between platform responsibility and scientist responsibility rather than letting it happen by accident. You can make the case for a platform investment that isn't obviously urgent yet, and you're comfortable saying no to a feature request that would compromise platform integrity.

The pay range for this role is:

145,000 - 180,000 USD per year (Hybrid (San Francisco, California, US))

Similar Jobs

More Information Technology Jobs

Find similar MLOps Engineer jobs: