Responsibilities
THE TEAM:
The Data Center GPU and Accelerated Processing Team delivers AMD products for Data Center, Machine Learning, and High-Performance Computing. The Platform Emulation team (PEMU) is an integral part of the Data Center GPU and Accelerated Processing group responsible for pre-silicon activities to support product bring-up and launch.
THE ROLE:
AMD is seeking a Platform Emulation Software Engineer to join our Data Center GPU organization. Our products support the rapidly scaling Data Center and High-Performance Compute infrastructure. You will be an integral member of the Platform Emulation team responsible for:
- Designing and implementing GPU applications in CUDA/C++ to enhance pre-silicon verification
- Writing tools to improve the efficiency of debugging hardware on emulation
- Building and executing GPU benchmarks/applications on the emulator
- Developing emulation infrastructure and tools
You will work alongside a team of innovative engineers to support the deployment of AMD’s Instinct ML products targeting Supercomputers and Data Center workloads.
THE PERSON:
The ideal candidate is an energetic, motivated engineer, excellent communication skills, and a solid technical foundation in GPU programming, and GPU architecture knowledge. This individual brings critical thinking, cross collaborate, and the ability to engage credibly with team members.
KEY RESPONSIBILITIES:
- Design and implement GPU applications in CUDA/HIP to enhance pre-silicon verification
- Develop tools and automation to improve the efficiency of debugging hardware on emulation
- Run and collect data for analysis on AMD’s high-end emulators and simulation models
- Develop scripts/tools to parse and analyze data from emulation runs
- Run and collect functional and performance data for AI/ML workloads
- Collaborate with senior engineers to support debug of hardware-related failures and performance issues observed on emulation
- Attend weekly meetings, provide status communication, and deliver technical presentations
SKILLS AND EXPERIENCE REQUIREMENTS:
Strong knowledge of computer hardware architecture (GPU/CPU, memory hierarchy, interconnects, caches) Excellent programming skills in:
EXPERIENCE OR UNDERSTANDING OF:
- Shared memory concurrent programming
- Relaxed memory models
- Cache coherency
- Working knowledge of Linux/Unix environments and shell scripting
- Knowledge of computer software architecture and boot flow (boot code, BIOS, device drivers, OS)
- Excellent oral and written communication skills
- Willingness to learn and think outside the box
NICE TO HAVE:
- Exposure to Verilog/SystemVerilog and waveform-based debug
- Experience with ML workloads and profiling/performance analysis
- Familiarity with APIs/platforms such as ROCm, OpenCL, OpenGL, Vulkan
WHAT YOU WILL LEARN:
- Data Center and Machine Learning use cases and workloads
- High-Performance Computing architecture and use cases
- How GPU-based software workloads interact with real hardware in a pre-silicon environment
- Pre-silicon development and testing technologies
- Hands-on emulation run and debug techniques
- Deep dives into performance analysis, code tracing, and debug on large-scale GPU platforms
ACADEMIC CREDENTIALS:
- Bachelor’s degree in Electrical Engineering, Computer Engineering, or equivalent experience.
- Master’s degree in a technical discipline or MBA is a plus.
LOCATION:
Markham, Canada preferred, open to Austin, TX.
This role is not eligible for visa sponsorship.
#LI-RW1
#LI-HYBRID
Qualifications
Benefits offered are described: AMD benefits at a glance.