5-7 years of experience in high performance ASIC architecture
Strong knowledge of silicon implementation and its constraints
Expertise in AI workloads, particularly low latency LLM inference
Familiarity with system-level bottlenecks related to memory and bandwidth
Hands-on experience with AI chip architecture and tapeouts preferred
Proven track record of end-to-end ownership of AI ASIC projects
Innovative thinking in proposing unconventional chip architectures
Responsibilities
Own the architecture of the XPU compute engine for A-0 and A-1 systems
Leverage existing 3rd party IP creatively to address bandwidth and memory challenges
Design semi-custom and custom XPU solutions for future generations
Analyze and optimize for low latency in AI workloads
Identify and mitigate workload and system-level bottlenecks
Collaborate with cross-functional teams to ensure architectural feasibility
Drive product delivery from concept through to implementation
Benefits
Opportunity to work on cutting-edge AI technology
Collaborative and innovative work environment
Potential for career growth in a rapidly evolving field
Access to advanced tools and resources for design and development
Engagement in projects with significant impact on future technology
Full Job Description
What We're Looking For:
We're looking for an XPU Architect to own the XPU compute engine architecture of the A-0 and A-1 systems. This will initially involve heavily leveraging existing 3rd party IP, but in a creative way to solve the bandwidth demand challenges and memory system interface challenges. Future generations of the compute XPU will be semi-custom and custom designs.
Ideal Background:
An ideal candidate will have:
Previous experience architecting and owning a high performance, high TDP ASIC (eg AI chips, network switch chips, high performance CPUs or similar)
Knowledge of both computer architecture but also silicon implementation, and understanding of the realistic constraints that silicon implementation places upon architectural options
Deep understanding of AI workloads, especially low latency LLM inference
Understanding of workload and system level bottlenecks, especially around inference, as it pertains to interconnect, memory capacity, bandwidth, latency, etc.
Strongly Preferred:
Hands-on experience with AI chip architecture and tapeouts
End to end ownership of at least one AI ASIC project
History of product delivery (regardless of ultimate commercial success of said product)
Experience with HBM-based AI chip architecture
Experience in proposing "weird" chip architectures (in any domain, not just AI!)