*Note: This position requires presence in our San Francisco or Bellevue office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
What You'll Do- Build and operate monitoring and alerting for cluster health - fabric, GPU, power/thermal, and job-level signals - to detect and respond to issues proactively
- Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible
- Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand
- Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely
- Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power - working closely with on-site deployment teams
- Participate in on-call rotations and lead incident response for cluster-level problems
- Contribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiency
You- 7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
- Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
- Strong understanding of Linux-based systems in a distributed environment
- Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
- Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling.
- Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
- Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
- Have excellent problem-solving and troubleshooting skills and an innate attention to detail
- Passion for continuous improvement and innovation
Nice to Have- Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
- Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
- Experience building and/or operating HPC resources.
- Depth in the NVIDIA hardware and firmware ecosystem
- Experience with data center power and thermal design
- Background in chaos engineering or similar reliability testing methodologies
- Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
We offer generous cash & equity compensation
Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use