*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
What You'll Do- Build and own fleet data systems that make Lambda's GPU datacenter hardware legible at scale - indexing, validating, and reconciling physical host state against intended logical configuration
- Lead the design and implementation of automation that orchestrates GPU cluster deployments from logical design import and racking through OS provisioning, validation, and customer hand-off
- Design systems that continuously validate consistency between intended and actual state across large-scale GPU environments, catching drift before it causes deployment failures or delays
- Design and implement ownership, global locking, readiness gating, and action safety systems that keep fleet operations coordinated across teams and tools
- Own Fleet Orc's technical contribution to new datacenter site bring-up, working cross-functionally with HPC Deployments, DC Ops, Network Engineering, and Core Infrastructure
- Monitor production health, hold SLAs, and drive resolution when things drift
You- Have 5+ years of engineering experience (degree not required)
- Are fluent in Python, Go, or similar - comfortable with APIs, distributed systems, and automation pipelines
- Can reason about PXE boot, firmware provisioning, IPMI/BMC interfaces, and network services like DNS and DHCP from first principles
- Have owned production systems with real SLAs
- Can lead technical design on medium-to-large features: take an ambiguous problem, write the doc, drive alignment, and ship
- Have worked cross-functionally and influenced technical decisions beyond your immediate team
- Have mentored peers and left systems - and teammates - better than you found them
Nice to Have- Experience in the machine learning or AI infrastructure industry
- Familiarity with datacenter physical infrastructure - racks, switches, InfiniBand fabric, power domains
- Background that blends software engineering with systems or infrastructure engineering
- Experience with network source-of-truth systems (NetBox or similar) or DCIM tooling
Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
We offer generous cash & equity compensation
Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use