THE ROLEAMD is seeking an AI Systems Engineer to help develop and optimize machine learning workloads on next-generation AMD AI accelerators. In this role, you will work at the intersection of hardware and software, designing high-performance ML operator kernels, optimizing dataflow pipelines, and enabling industry-leading AI inference performance across AMD NPU and GPU platforms.
You will collaborate closely with compiler, runtime, silicon, and architecture teams while helping bring cutting-edge AI technologies from concept to production. This role offers full-stack visibility from kernel development and model optimization through hardware validation and silicon bring-up. If you are passionate about AI systems, accelerator architectures, and solving complex performance challenges, this is an opportunity to make a significant impact on products deployed in millions of devices worldwide.
THE PERSONThe ideal candidate is a systems-minded engineer who enjoys tackling complex performance and optimization challenges at the hardware-software boundary. You are naturally curious, thrive in collaborative environments, and are comfortable working across multiple technical domains to debug, analyze, and improve system behavior.
You have a strong foundation in computer architecture and machine learning systems, enjoy working with cross-functional teams, and can translate technical insights into scalable solutions that improve performance, functionality, and product quality.
KEY RESPONSIBILITIES- Develop and optimize machine learning operator kernels and dataflow libraries for AMD AI accelerators.
- Profile workloads, identify performance bottlenecks, and drive software and system-level optimizations.
- Enable and validate ML models within production inference frameworks and runtime environments.
- Collaborate with compiler, runtime, architecture, and silicon teams to deliver high-performance AI solutions.
- Debug and resolve issues spanning kernel implementation, runtime integration, model accuracy, and hardware bring-up.
- Contribute to hardware-software co-design efforts by evaluating architectural tradeoffs and influencing future accelerator technologies.
- Drive innovation in performance methodologies, benchmarking, tooling, and AI system optimization.
PREFERRED EXPERIENCE- Strong software development experience using C/C++ and Python.
- Experience with parallel programming, multithreaded applications, and performance optimization.
- Knowledge of machine learning inference workloads and common operators such as GEMM, convolution, attention, and softmax.
- Familiarity with AI frameworks and runtimes such as PyTorch, ONNX Runtime, ROCm, or similar technologies.
- Understanding of computer architecture, memory hierarchies, cache behavior, and accelerator programming models.
- Experience developing software for GPUs, NPUs, AI accelerators, or other high-performance computing platforms.
- Experience using development, debugging, profiling, and source control tools in Linux environments.
- Familiarity with MLIR, LLVM, compiler technologies, or related software stacks.
- Exposure to quantization techniques, including INT8, FP8, FP16, or BF16 optimization.
- Knowledge of dataflow architectures, systolic arrays, or custom accelerator designs.
- Publications, patents, or demonstrated technical contributions in machine learning systems, computer architecture, or related fields.
ACADEMIC CREDENTIALS- Master's or PhD in Computer Engineering, Electrical Engineering, Computer Science, or a related technical field preferred.
LOCATIONSan Jose, CA
This role is not eligible for visa sponsorship.#LI-DR2#LI-HYBRIDBenefits offered are described: AMD benefits at a glance.