MLOps / ML Platform Engineer

SumerSports LLC

$170K — $200K *
US-AnywhereRemote in United States
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years of experience in ML platform, DevOps, or infrastructure engineering.
  • Deep knowledge of Kubernetes, CI/CD, containers, and cloud services (AWS, GCP, or Azure).
  • Hands-on experience managing GPU clusters and training/inference pipelines.
  • Familiarity with data orchestration and formats like Delta, Parquet, and Spark.
  • Proven ability to ship and operate production ML systems with SLOs.
  • Strong Python coding skills and experience with infrastructure automation.
  • Experience in optimizing observability and costs at scale.

Responsibilities

  • Design and operate ML infrastructure for data management and high-throughput workflows.
  • Build scalable, reproducible training and evaluation pipelines with versioning.
  • Optimize compute usage and costs by managing GPU/CPU workloads and clusters.
  • Serve production models via APIs with low latency and robust rollback processes.
  • Ensure reliability by tracking key performance indicators like latency and cost.
  • Automate deployment and security processes through CI/CD and infrastructure as code.
  • Collaborate with cross-functional teams to streamline model delivery from experimentation to production.
  • Create documentation and tools for repeatable and fast ML workflows.

Benefits

  • Competitive salary and bonus plan.
  • Comprehensive health insurance plan.
  • Retirement savings plan (401k) with company match.
  • Remote working environment.
  • Flexible, unlimited time off policy.
  • Generous holiday schedule, including 13 holidays and additional time off after the Super Bowl.
Full Job Description
Responsibilities

  • Design and operate ML infrastructure: Manage data, training, serving, and inference systems for high-throughput model workflows.
  • Build scalable pipelines: Implement reproducible training and evaluation pipelines with versioning, scheduling, and artifact tracking.
  • Optimize compute and cost: Tune GPU and CPU workloads, manage clusters, and drive efficiency via rightsizing, spot scheduling, and caching.
  • Serve models in production: Operate APIs for low-latency inference with autoscaling, blue-green or canary rollouts, and rollback safety.
  • Ensure reliability and observability: Define and own SLOs; instrument pipelines and services to track latency, cost, drift, and data quality.
  • Secure and automate: Manage IAM, secrets, and container security; automate deployment pipelines via CI/CD and infrastructure as code.
  • Collaborate cross-functionally: Partner with research scientists and AI engineers to deliver models from experiment to production with minimal friction.
  • Document and enable: Build templates, runbooks, and internal tooling that make ML workflows repeatable, safe, and fast.

Qualifications

  • 4+ years of experience in ML platform, DevOps, or infrastructure engineering.
  • Deep knowledge of Kubernetes, CI/CD, containers, and cloud infrastructure (AWS, GCP, or Azure).
  • Hands-on experience managing GPU clusters and training/inference pipelines.
  • Familiarity with data orchestration and storage formats (Delta, Parquet, Polars, Spark).
  • Proven ability to ship and operate production ML systems with SLOs.
  • Strong Python skills and comfort with infrastructure as code and automation.
  • Experience with observability and cost optimization at scale.

Nice to Have

  • Experience with real-time or low-latency model serving (REST, gRPC).
  • Exposure to model registry and promotion workflows.
  • Familiarity with data quality, lineage, and curation pipelines.
  • Background in sports analytics or other high-volume data domains.
  • Experience integrating LLM workflows or evaluation pipelines.

Benefits

  • Competitive Salary and Bonus Plan
  • Comprehensive health insurance plan
  • Retirement savings plan (401k) with company match
  • Remote working environment
  • A flexible, unlimited time off policy
  • Generous paid holiday schedule - 13 in total including Monday after the Super Bowl


SumerSports is committed to fair and equitable compensation practices.

Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location. This may be different in other locations due to differences in the cost of labor.

The total compensation package for this position may also include annual performance bonus, benefits and/or other applicable incentive compensation plans.

The pay range for this role is:

170,000 - 200,000 USD per year (Remote)

Similar Jobs

More Enterprise Technology Jobs

Find similar MLOps / ML Platform Engineer jobs: