*Note: This position requires presence in our San Jose, San Francisco, or Bellevue office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
What You'll Do- Architect and define scalable compute platforms optimized for AI/ML, simulation, and high-throughput workloads.
- Develop compute system standards and design patterns to ensure consistency, performance, and maintainability across infrastructure.
- Evaluate emerging CPU, GPU, and accelerator technologies, owning architectural tradeoff decisions that impact compute density, power, cooling, and total cost.
- Collaborate with product and engineering teams to map workload requirements to compute platform capabilities across bare metal and cloud deployments.
- Experience converting ambiguous business or customer needs into measurable platform requirements, technical specifications, acceptance criteria, and architecture decisions.
- Define compute platform roadmaps and architectural reference designs that guide hardware selection, firmware baselines, rack-level, and cluster design.
- Act as a technical lead during new platform introductions, guiding validation and performance characterization efforts.
- Mentor systems engineers and cross-functional stakeholders on compute performance tuning, sizing, and architectural decisions.
You- Proven experience (7+ years) architecting large-scale 10k-100k+ GPU HPC or cloud compute platforms.
- Deep knowledge of CPU/GPU architectures, memory hierarchies, and accelerator topologies.
- Experience designing systems around high-bandwidth, low-latency fabrics (NVLink, InfiniBand, and RoCE).
- Strong understanding of system performance tuning, resource scheduling, thermal and power optimization, and compute lifecycle management.
- Comfortable working across hardware and software boundaries, especially at the intersection of compute architecture, OS behavior, and orchestration layers.
- Skilled at balancing architectural tradeoffs for density, power efficiency, cooling, and performance.
- Strong analytical and communication skills, with a track record of influencing technical strategy across teams.
- Strong ownership and can do attitude, self-starter who feels comfortable working in ambiguity.
Nice to Have- Hands-on experience with AI/ML workloads and their compute performance characteristics.
- Familiarity with orchestration tools used in HPC. (Slurm, Kubernetes, etc)
- Experience with virtualization technologies, specifically GPU virtualization.
- Exposure to hardware validation, vendor collaboration, and long-term OEM roadmap alignment.
- Background in compute telemetry, real-time performance profiling, or large-scale A/B infrastructure testing.
Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
- Founded in 2012, with 500+ employees, and growing fast
- Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove
- We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG
- Our values are publicly available: https://lambda.ai/careers
- We offer generous cash & equity compensation
- Health, dental, and vision coverage for you and your dependents
- Wellness and commuter stipends for select roles
- 401k Plan with 2% company match (USA employees)
- Flexible paid time off plan that we all actually use