5-7 years of experience in high-performance ASIC architecture, particularly in AI or networking domains.
Strong knowledge of computer architecture and silicon implementation constraints.
Expertise in AI workloads, focusing on low latency LLM inference.
Familiarity with system-level bottlenecks related to memory and interconnect performance.
Hands-on experience with AI chip architecture and tapeouts is a plus.
Responsibilities
Own the architecture of the XPU compute engine for A-0 and A-1 systems.
Leverage existing 3rd party IP creatively to address bandwidth and memory challenges.
Design semi-custom and custom XPU architectures for future generations.
Analyze and optimize for low latency and high throughput in AI workloads.
Identify and mitigate system-level bottlenecks in inference processes.
Benefits
Opportunity to lead innovative projects in cutting-edge AI technology.
Collaborative work environment with a focus on creativity and problem-solving.
Access to advanced tools and resources for chip design and architecture.
Potential for professional growth in a rapidly evolving field.
Full Job Description
What We're Looking For:
We're looking for an XPU Architect to own the XPU compute engine architecture of the A-0 and A-1 systems. This will initially involve heavily leveraging existing 3rd party IP, but in a creative way to solve the bandwidth demand challenges and memory system interface challenges. Future generations of the compute XPU will be semi-custom and custom designs.
Ideal Background:
An ideal candidate will have:
Previous experience architecting and owning a high performance, high TDP ASIC (eg AI chips, network switch chips, high performance CPUs or similar)
Knowledge of both computer architecture but also silicon implementation, and understanding of the realistic constraints that silicon implementation places upon architectural options
Deep understanding of AI workloads, especially low latency LLM inference
Understanding of workload and system level bottlenecks, especially around inference, as it pertains to interconnect, memory capacity, bandwidth, latency, etc.
Strongly Preferred:
Hands-on experience with AI chip architecture and tapeouts
End to end ownership of at least one AI ASIC project
History of product delivery (regardless of ultimate commercial success of said product)
Experience with HBM-based AI chip architecture
Experience in proposing "weird" chip architectures (in any domain, not just AI!)