5-7 years of experience in high-performance ASIC architecture, particularly in AI or networking domains.
Strong knowledge of computer architecture and silicon implementation constraints.
Expertise in AI workloads, focusing on low latency LLM inference.
Familiarity with system-level bottlenecks related to memory and interconnect performance.
Hands-on experience with AI chip architecture and tapeouts is a plus.
Proven track record of end-to-end ownership of AI ASIC projects.
Innovative thinking in proposing unconventional chip architectures.
Responsibilities
Own the architecture of the XPU compute engine for A-0 and A-1 systems.
Leverage existing third-party IP creatively to address bandwidth and memory challenges.
Design semi-custom and custom XPU solutions for future generations.
Analyze and optimize for low latency in AI workloads.
Identify and mitigate system-level bottlenecks in inference processes.
Collaborate with cross-functional teams to ensure architectural feasibility.
Drive product delivery timelines and milestones for AI ASIC projects.
Benefits
Opportunity to work on cutting-edge AI technology.
Collaborative and innovative work environment.
Potential for career growth in a rapidly evolving field.
Access to advanced tools and resources for design and development.
Full Job Description
What We're Looking For:
We're looking for an XPU Architect to own the XPU compute engine architecture of the A-0 and A-1 systems. This will initially involve heavily leveraging existing 3rd party IP, but in a creative way to solve the bandwidth demand challenges and memory system interface challenges. Future generations of the compute XPU will be semi-custom and custom designs.
Ideal Background:
An ideal candidate will have:
Previous experience architecting and owning a high performance, high TDP ASIC (eg AI chips, network switch chips, high performance CPUs or similar)
Knowledge of both computer architecture but also silicon implementation, and understanding of the realistic constraints that silicon implementation places upon architectural options
Deep understanding of AI workloads, especially low latency LLM inference
Understanding of workload and system level bottlenecks, especially around inference, as it pertains to interconnect, memory capacity, bandwidth, latency, etc.
Strongly Preferred:
Hands-on experience with AI chip architecture and tapeouts
End to end ownership of at least one AI ASIC project
History of product delivery (regardless of ultimate commercial success of said product)
Experience with HBM-based AI chip architecture
Experience in proposing "weird" chip architectures (in any domain, not just AI!)