About the RoleAs a Staff Software Engineer for the Compute pillar, you will play a critical role in defining the technical vision for Lambda's next-generation GPU and CPU host instance lifecycle and compute control plane. This role bridges the gap between high-level distributed systems and low-level semiconductor architecture to enable seamless, reliable cloud provisioning and lifecycle management of a heterogeneous compute platform at a massive scale. You will provide hands-on technical leadership that will guide development of a resilient compute control plane utilizing durable execution concepts and deep/unique hardware integration.
The position requires a deep understanding of the entire stack, from BIOS/firmware (UEFI), Linux kernel internals, modern DPU capabilities, distributed systems, cradle-to-grave system lifecycle management, to large-scale cloud-service provider (CSP) operations. You will drive high-impact, cross-functional initiatives, leading the work of multiple engineers to deliver enterprise-grade SLAs for the world's leading AI researchers.
What You'll DoWe are seeking an engineer with extensive experience in cloud infrastructure to build and optimize GPU-first compute systems. In this role, you will be responsible for:
- Designing and implementing a highly available and reliable GPU and CPU "host and instance lifecycle" control plane.
- Guide technical decisions involving semiconductor architecture, BIOS/Firmware settings, system boot methodologies, and DPU utilization to optimize host capabilities, performance and reliability.
- Guide design of compute platform multi-tenant security model
- Provide technical leadership and mentorship for senior engineers across several teams to execute on complex infrastructure roadmaps and technical strategy.
- Collaborate with product and data center organizations to translate customer requirements into scalable infrastructure capabilities.
- Work with customers on translating vague customer technical requirements into concrete engineering deliverables.
- Set engineering standards and lead design reviews for mission-critical cloud software at scale.
Who You are- 10+ years of experience working on compute control plane distributed systems used for deploying and lifecycle managing heterogeneous compute platforms into data-centers, built for resilience at scale.
- Deep expertise in durable execution models and distributed systems used in cloud-service provisioning.
- Basic knowledge of software defined networking fundamentals that informs secure, multi-tenant distributed systems.
- Proven track record of leading large-scale semi-conductor hardware enablement and deployment initiatives.
- Proven experience in deploying net-new data-centers into a global compute platform (not just working in existing data-centers).
- Proficiency in one of more of the following programming languages: C/C++, Rust, Python, Go.
Nice to Have- Knowledge of Nvidia's AI Factory architectural components (including GPU hosts, CPU hosts, SuperNICs (ConnectX and Bluefield DPUs , and switches).
- Knowledge of Nvidia's AI Factory software offerings (like DOCA, DOCA SNAP, CUDA, et al.)
- Knowledge of Linux kernel internals, device drivers, and virtualization technologies (KVM, QEMU), kernel bypass technologies (like SR-IOV, DPDK, SPDK).
- Experience with Cloud Service Provider Kubernetes offerings.
- Knowledge of high-performance networking (InfiniBand, RoCE) and storage protocols (NVMe-oF).
Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
About Lambda- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use