ML Platform Engineer

UniversalAGI

• $150K — $180K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong software engineering skills (clean code, debugging, reliability, reproducibility).
  • Hands-on experience building/operating infrastructure for ML/compute-heavy workflows: pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring.
  • Olympic athlete mindset: obsession with measurable improvement.
  • Resourcefulness: knowing when to do quick fixes versus robust solutions.
  • Ownership: being accountable for work end-to-end.

Responsibilities

  • Build and operate scalable infrastructure for data generation and simulation workflows.
  • Build reproducible pipelines for training/fine-tuning and benchmarking.
  • Own cost/performance tradeoffs in compute and storage.
  • Lead deployments of stack into customer environments (AWS/GCP/Azure + hybrid).
  • Build robust deployment patterns: CI/CD, monitoring, incident response.
  • Partner with customers to ensure reliability and compliance under constraints.

Benefits

  • Competitive compensation and equity.
  • Paid health, dental, and vision benefits.
  • 401(k) plan offering.
  • Flexible vacation policy.
  • Team building and fun activities provided.
  • Great scope for ownership and impact.
  • Stipend for AI tools.
  • Monthly commute stipend.
  • Monthly wellness/fitness stipend.
  • Daily meals provided by the company.
  • Immigration support available.
Full Job Description
San Francisco | 5 Days Onsite

Location: Onsite in San Francisco

Compensation: Competitive Salary + Equity

About the Role

UniversalAGI is hiring a ML Platform Engineer to build and own the execution platform powering our research and customer deployments: data generation + simulation orchestration + training/fine-tuning infrastructure + benchmarking pipelines + production deployments in customer environments.

You'll work closely with the CEO and founding team to turn research into repeatable, scalable, reliable systems - internally and in customer infrastructure. This is a "ship outcomes" role: your work directly determines how fast we can iterate, how reproducible our results are, and how reliably we deliver in production.

What You'll Do

Build the foundation platform (internal):
  • Build and operate scalable infrastructure for data generation and simulation workflows (job orchestration, scheduling, queues, retries, observability).
  • Build reproducible pipelines for training/fine-tuning and benchmarking (artifact/version management, experiment tracking, dataset lineage).
  • Own cost/performance tradeoffs across compute, storage, networking, and runtime efficiency.

Deploy to customers (external):
  • Lead deployments of our stack into customer cloud/on-prem environments (AWS/GCP/Azure + hybrid), including secure networking, permissions, and data movement.
  • Build robust deployment patterns: environment provisioning, CI/CD, rollbacks, monitoring, and incident response.
  • Partner with customers to ensure reliability and repeatability under real-world constraints (security, compliance, infra limits, data governance).


Qualifications
  • Strong software engineering skills (clean code, debugging, reliability, reproducibility).
  • Hands-on experience building/operating infrastructure for ML/compute-heavy workflows: pipelines, job orchestration, GPU compute, storage, CI/CD, monitoring.
  • Olympic athlete mindset: You have high standards for yourself and are obsessed with measurable improvement on the metrics you are delivering to customers.
  • Resourcefulness: you know when to do the "quick & correct" fix vs. when to invest in a robust solution, and you can justify the tradeoff with impact/
  • Ownership: Comfortable owning work end-to-end and being accountable for measurable outcomes.


Bonus Qualifications
  • Experience with workflow orchestration (e.g., Ray, Kubernetes, Slurm).
  • Experience with GPU infrastructure and distributed training systems.
  • Experience building evaluation/benchmarking frameworks with strong reproducibility guarantees.
  • Experience deploying into regulated / security-sensitive environments (gov/defense/enterprise).
  • Experience with simulation/HPC pipelines (CFD, meshing, batch workloads) is a plus but not required.
  • Experience in an FDE-style / delivery execution role (or similar "ship results fast" environments).


Cultural Fit
  • Technical Respect: Ability to earn respect through hands-on technical contribution
  • Intensity: Thrives in our unusually intense culture - willing to grind when needed
  • Customer Obsession: Passionate about solving real customer problems, not just cool tech
  • Deep Work: Values long, uninterrupted periods of focused work over meetings
  • High Availability: Ready to be deeply involved whenever critical issues arise
  • Communication: Can translate complex technical concepts to customers and team
  • Growth Mindset: Embraces the compounding returns of intelligence and continuous learning
  • Startup Mindset: Comfortable with ambiguity, rapid change, and wearing multiple hats
  • Work Ethic: Willing to put in the extra hours when needed to hit critical milestones
  • Team Player: Collaborative approach with low ego and high accountability


What We Offer
  • Opportunity to shape the technical foundation of a rapidly growing foundational AI company.
  • Work on cutting-edge industrial AI problems with immediate real-world impact.
  • Direct collaboration with the founder & CEO and ability to influence company strategy
  • Competitive compensation with significant equity upside.
  • In-person first culture - 5 days a week in office with a team that values face-to-face collaboration.
  • Access to world-class investors and advisors in the AI space.


Benefits

We provide great benefits, including:
  • Competitive compensation and equity.
  • Competitive health, dental, vision benefits paid by the company.
  • 401(k) plan offering.
  • Flexible vacation.
  • Team Building & Fun Activities.
  • Great scope, ownership and impact.
  • AI tools stipend.
  • Monthly commute stipend.
  • Monthly wellness / fitness stipend.
  • Daily office lunch & dinner covered by the company.
  • Immigration support.

Similar Jobs

More Jobs at UniversalAGI

More Enterprise Technology Jobs

Find similar ML Platform Engineer jobs: