As a Software Engineer for the ML Ops team, you will help own the production runtime for Phare's ML stack - deploying, serving, and scaling models across inference endpoints and batch/streaming workflows.
Every day you will build progressive delivery pipelines with automated rollouts and rollbacks, manage SLOs for latency and availability, and instrument end-to-end observability (metrics, logs, traces, drift, regression). You'll harden the platform with Terraform, Kubernetes, and CI/CD, ensuring reproducible, auditable ML releases.
To thrive in this role, you must have a background in operating ML systems at scale, where uptime and feedback loops matter as much as accuracy.
We prefer this role to be Hybrid (3 days onsite) at our SOHO Office Hub in New York City.
Required Skills: - Production ML: You've deployed and operated models running on GPUs in production - APIs and batch/streaming inference
- Platform engineering: Strong with Docker/Kubernetes, IaaC (e.g., Terraform), and CI/CD for services and model artifacts; you maintain environment parity, reproducible releases, and robust model/experiment versioning with data lineage
- System Reliability: You use progressive delivery with automated rollouts/rollbacks, and you build end-to-end observability (metrics, logs, traces, and model telemetry for drift/regression) plus actionable alerting, runbooks, and incident response
- Post-training lifecycles: You manage model registries and stage gates, design scheduled or event-driven retraining when appropriate, and enforce RBAC, secrets management, encryption, and audit logs
Bonus if you have: Experience in regulated environments (e.g., Healthcare, Finance)
Role Leveling: We are looking for candidates at various levels, ranging from Level 2 to Staff
- L2: Independently delivers a complete end-to-end project, owning design, implementation, and delivery of scoped work (Compensation Range for L2: $140k - $200k)
- L3: Leads delivery of larger projects, handling increased technical complexity and ambiguity, providing light guidance to L2s on shared work (Compensation Range for L3: $140k - $253k)
- Senior: Team Lead responsible for managing a portfolio of projects that contribute to major technical initiatives. (Compensation Range for Senior: $140k - $294k)
- Staff: Impact at the organizational level. Responsible for leading multiple teams or multiple broad initiatives simultaneously, ensuring that high-level technical goals are met across the entire organization. (Compensation Range for Staff: $140k - $350k)
Benefits - Top-of-market compensation, including bonus that starts at 10%
- Comprehensive health benefits
- Inspiring, brilliant, mission-driven teammates
For this US-based position, the base pay range is $140,000.00 - $270,000.00 per year . Individual pay is determined by role, level, location, job-related skills, experience, and relevant education or training.
This job is eligible to participate in our annual bonus plan at a target of 10.00%