Unity Technologies

Senior Machine Learning Engineer, ML Infrastructure- Online

Unity Technologies$187K — $243K *
US-AnywhereRemote in Seattle, WA
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in production ML inference systems
  • Strong programming skills in Python, particularly in high-scale services
  • Deep knowledge of model serving frameworks (e.g., TensorFlow Serving, NVIDIA Triton)
  • Experience with distributed systems and Kubernetes for service reliability
  • Proven ability to influence architectural direction across teams
  • Experience in optimizing ML model performance and inference workloads
  • Strong understanding of safe model rollout strategies and service observability.

Responsibilities

  • Design and operate online inference infrastructure for ML models
  • Develop infrastructure for distributed training workflows
  • Integrate ML pipelines with orchestration systems to ensure reliability
  • Optimize model performance through various technical improvements
  • Enhance observability of ML systems to monitor performance metrics
  • Collaborate with ML engineers to support model iteration and efficiency
  • Lead architectural improvements for a more robust ML platform.

Benefits

  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Mental Health and Wellbeing programs and support
Full Job Description
The Role

We are seeking a Senior ML engineer to design and evolve Unity Vector's online model inference platform. This role focuses on building reliable infrastructure for serving machine learning models in production, optimizing inference performance, and enabling safe, efficient experimentation across high-traffic online systems.

You will work closely with ML engineers, platform teams, and product stakeholders to ensure models can be deployed, scaled, monitored, and iterated on efficiently. You will play a key role in shaping how models are packaged, served, validated, monitored, and optimized in production environments.

This role requires strong systems thinking, deep experience with production ML infrastructure, and the ability to drive architectural improvements across teams.

What you'll be doing
  • Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or similar distributed serving frameworks.
  • Develop infrastructure that supports distributed training workflows using technologies such as Pytorch, Ray Data, and Ray Train, etc.
  • Integrate ML pipelines with workflow orchestration systems (e.g., Flyte, Airflow, or similar) to enable reliable multi-stage training workflows
  • Optimize model performance through model compilation, GPU/CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
  • Improve observability of ML systems through latency, throughput, error-rate, cost, saturation, and model-health monitoring.
  • Partner closely with ML engineers to support faster model iteration while maintaining production safety, scalability, and cost efficiency.
  • Improve the reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
  • Lead architectural improvements that make the online ML platform more robust, user-friendly, scalable, and cost-efficient.


What we're looking for
  • Experience building and operating production-grade online ML inference systems, such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
  • Experience with model serving frameworks such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
  • Experience optimizing inference workloads using techniques such as dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
  • Strong experience with distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
  • Strong programming skills in Python, with practical experience working on production ML systems and high-scale services.
  • Experience with PyTorch and modern model deployment workflows, including model packaging, validation, and serving lifecycle management.
  • Experience designing infrastructure for safe model rollout, canary testing, A/B experimentation, and automated rollback.
  • Strong systems thinking, with the ability to reason about latency, throughput, reliability, scalability, and cost tradeoffs in online systems.
  • Proven ability to lead technical direction and influence architectural decisions across teams without formal authority.


Additional information
  • Relocation support is not available for this position
  • Work visa/immigration sponsorship is not available for this position
  • $187,200$-$243,300


This range reflects the anticipated base salary for this position. Beyond base salary, this role may be eligible for equity awards and participation in our company incentive plans (such as annual discretionary bonuses or sales commissions). The final offer amount will depend on several factors, including geographic location and the candidate's relevant experience, professional background, and skill set.

Benefits

At Unity, we want our team members to thrive. We offer a wide range of benefits designed to support well-being and work-life balance.

Please note: Benefits eligibility, specific offerings, and coverage vary based on the country and employment status.

While specific benefits vary, here are some of the ways we strive to take care of our eligible team members globally: Comprehensive health, life, and disability insurance | Commute subsidy | Employee stock ownership | Competitive retirement/pension plans | Generous vacation and personal days | Support for new parents through leave and family-care programs | Office food snacks | Mental Health and Wellbeing programs and support | Employee Resource Groups | Global Employee Assistance Program | Training and development programs | Volunteering and donation matching program

About Unity Technologies

Unity Technologies is a software company that provides a platform for creating and operating interactive, real-time 3D content. The company's platform is used by game developers, architects, automotive designers, filmmakers, and other creators to build and distribute interactive experiences. Unity Technologies was founded in 2004 and is headquartered in San Francisco, California. The company has offices in North America, Europe, and Asia.
Learn more about Unity Technologies
Size
4,000 employees
Industry
Founded
2004

Similar Jobs

More Jobs at Unity Technologies

More Information Technology Jobs

Find similar Senior Machine Learning Engineer, ML Infrastructure- Online jobs: