*Note: This position requires presence in our San Francisco or San Jose office location 4 days per week; Lambda's designated work from home day is currently Tuesday.
Architecture leads the design, development, and execution of next-generation GPU clusters. Our team gets first access to the latest technologies and is responsible for translating emerging platform capabilities into scalable, production-ready infrastructure. Our scope spans from rack and pod level system arrangement that maximizes the data center power footprint, to right-sizing compute fabrics that meet the most demanding customer workloads.
We work at the intersection of GPU systems, high-performance networking, power and cooling, data center design, and customer-scale AI infrastructure.
We're looking for a Senior HPC Systems Architect with extensive experience designing, developing, and testing large-scale high-performance computing (HPC) infrastructures. This strategic role focuses on crafting cutting-edge liquid-cooled HPC solutions.
What You'll Do- Design and architect advanced HPC systems optimized for large-scale computational workloads and AI applications.
- Collaborate with internal teams and stakeholders to define system requirements and performance goals.
- Develop comprehensive testing frameworks to rigorously assess system performance, scalability, and reliability.
- Evaluate emerging technologies and architectural approaches to continuously enhance infrastructure capabilities.
- Create detailed architectural plans, documentation, and blueprints to guide implementation teams.
- Provide technical leadership and mentoring to engineering teams, fostering best practices in HPC architecture.
You- 8+ years of experience designing and architecting large-scale HPC and distributed computing systems.
- Expert-level knowledge of HPC hardware including GPU clusters, compute nodes, high-speed networking (InfiniBand, Ethernet), and distributed storage.
- Hands-on experience with direct-to-chip liquid cooling systems.
- Proven expertise in creating robust performance benchmarks, capacity planning, and system validation.
- Exceptional skills in system architecture, design documentation, and technical specifications.
- Ability to work collaboratively across teams, ensuring alignment of technical solutions with business objectives.
- Self-motivated, strategic thinker with strong analytical and problem-solving capabilities.
Nice to Have- Prior experience with AI infrastructure design.
- Familiarity with cloud computing environments and hybrid cloud HPC architectures.
- Knowledge of automation and orchestration tools (Ansible, Terraform, Kubernetes).
Salary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.