Senior Staff AI Accelerator Performance Architect

Cerebras Systems

$175K — $275K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years in performance analysis or architecture for high-performance computing systems.
  • Strong hardware architecture knowledge from compiler, kernel, or performance work.
  • Experience with analytical or simulation-based performance models using Python, C++ or similar.
  • Solid grasp of processor architecture, memory systems, and hardware constraints.
  • Ability to link kernel behavior with application/system performance.
  • Experience in profiling workloads and validating performance hypotheses.
  • Clear communication of modeling assumptions and recommendations.

Responsibilities

  • Own and develop performance models for next-gen accelerator architectures.
  • Build analytical and simulation-based models across workloads and generations.
  • Analyze AI workloads to identify resource utilization and bottlenecks.
  • Evaluate architectural features for performance improvements.
  • Collaborate with cross-functional teams on optimizations and mappings.
  • Create workload projections and performance analyses based on solid assumptions.
  • Define success criteria and performance targets for future products.

Benefits

  • Opportunity to influence hardware and software roadmaps.
  • Engage in cutting-edge AI and performance architecture projects.
  • Collaborative work environment with cross-functional teams.
  • Potential for professional growth in a rapidly evolving tech landscape.
Full Job Description
Senior Staff AI Accelerator Performance Architect

Wafer-scale computing creates a distinctive architecture space in which compute placement, memory capacity and bandwidth, communication, kernel execution and system-level behavior must be understood together.

We are looking for a performance architect to guide the evolution of our next-generation AI systems. You will connect real workloads to architectural behavior, identify the bottlenecks that matter, quantify potential improvements and influence hardware and software roadmaps through rigorous performance analysis.

This role is ideal for someone with deep knowledge of hardware architecture, developed through hardware, compiler, kernel or system-performance work, who enjoys operating at the intersection of applications, kernels, architecture and system performance.

What You'll Do
  • Own and evolve performance models and modeling methodologies for next-generation accelerator and system architectures.
  • Build and extend analytical, simulation-based or trace-driven models across workloads, architectural features and product generations.
  • Analyze important AI workloads, from individual kernels through end-to-end inference and training execution, to determine where time, bandwidth, compute and capacity are spent.
  • Identify hardware and software bottlenecks and quantify opportunities to improve latency, throughput, utilization and energy efficiency.
  • Evaluate proposed architectural features and determine their expected performance return across representative workloads.
  • Study how models and kernels map onto the underlying compute, memory and communication architecture.
  • Partner with architecture, compiler, kernel, runtime and systems teams to evaluate alternative mappings and optimizations.
  • Develop workload projections and competitive performance analyses grounded in transparent assumptions.
  • Create concise recommendations that translate complex performance results into architectural and product decisions.
  • Improve modeling methodology, validation and correlation with RTL, emulation and silicon measurements.
  • Help define representative workloads, performance targets and success criteria for future products.
What We're Looking For
  • 7+ years of experience in performance analysis, performance modeling or architecture exploration for CPUs, GPUs, AI accelerators or other high-performance computing systems.
  • Strong understanding of hardware architecture developed through hardware, compiler, kernel, runtime or system-performance work.
  • Experience developing analytical, simulation-based or trace-driven performance models using Python, C++ or similar environments.
  • Solid understanding of processor architecture, memory systems, interconnects, parallel execution and hardware resource constraints.
  • Ability to move between kernel-level behavior and end-to-end application or system performance.
  • Experience profiling workloads, forming performance hypotheses and validating them with quantitative evidence.
  • Understanding of how software mapping and programmability affect realized hardware performance.
  • Ability to communicate modeling assumptions, uncertainty, bottlenecks and recommendations clearly.
  • MS or PhD in Electrical Engineering, Computer Engineering, Computer Science or equivalent practical experience.
Particularly Relevant Experience
  • Performance analysis of transformer inference or training workloads.
  • Attention, GEMM/GEMV, collective communication, mixture-of-experts, quantization or memory-capacity-constrained execution.
  • Kernel optimization, compiler performance, runtime scheduling or distributed accelerator systems.
  • Model validation using RTL simulation, emulation, FPGA prototypes or silicon measurements.
  • Competitive analysis of AI accelerators and large-scale AI systems.

Role Focus

This is a performance and architecture role, not a production RTL-design position. You should be comfortable reasoning about microarchitecture and working with architecture, RTL and physical-design teams, but you will not be expected to own detailed microarchitecture specifications, production RTL implementation, synthesis closure or physical design.

This role evaluates architectural features and recommends improvements; the AI Accelerator Architect owns the detailed feature definition and implementation-ready microarchitecture specification.

Your primary deliverables are trusted models, workload insights, feature ROI and architectural recommendations.

The base salary range for this position is $175,000 to $275,000 annually. Actual compensation may include bonus and equity, and will be determined based on factors such as experience, skills, and qualifications.

Similar Jobs

More Jobs at Cerebras Systems

More Information Technology Jobs

Find similar Senior Staff AI Accelerator Performance Architect jobs: