Advanced Micro Devices, Inc

Principal Software Quality Engineer - GPU & Machine Learning

Advanced Micro Devices, Inc$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of software engineering experience in validation, SDET, or quality engineering, with a focus on complex systems validation.
  • Proficient in Python for test automation; strong C++ skills for debugging and extending production code.
  • Expertise in GPU software stacks (e.g., ROCm, CUDA) and AI/ML frameworks (e.g., PyTorch, TensorFlow).
  • Experience with validation on multi-GPU, multi-node server platforms, including stress and fault injection testing.
  • Familiarity with the latest validation practices, including agile quality engineering and automated testing.
  • Track record of contributing to open-source projects, particularly in validation and CI pipelines.
  • Advanced academic credentials in Computer Science, Computer Engineering, or related fields.

Responsibilities

  • Own the end-to-end validation architecture for ROCm across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases.
  • Architect and evolve the test infrastructure for distributed testing and CI integration.
  • Champion and implement modern quality engineering methodologies within the team.
  • Set standards for GitHub workflows, including PR gating and issue management.
  • Lead troubleshooting for complex multi-component failures and enhance test coverage.
  • Influence product roadmaps with validation readiness for emerging technologies.

Benefits

  • Hybrid work environment with flexible working hours.
  • Opportunity to work with cutting-edge technology in AI and GPU computing.
  • Mentorship roles that foster professional growth and skills development.
  • Engagement with strategic customers and participation in open-source initiatives.
  • Access to advanced hardware and resources for testing and development.
Full Job Description
THE ROLE:

We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm software validation across compute workloads and server-class systems. In this individual-contributor leadership role, you will define how AMD proves ROCm is ready to ship - from unit and component testing, through full-stack workload validation, to multi-node system-level qualification on AMD Instinct™ GPU platforms. You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

THE PERSON:

You will set the technical direction for validation strategy, build and evolve the test infrastructure that gates every ROCm release, and personally drive the hardest debugging, characterization, and qualification problems. Your work directly determines the quality bar experienced by hyperscalers, OEMs, sovereign-AI customers, and the open-source community running ROCm in production.

KEY RESPONSIBILITIES:

  • Own the end-to-end validation architecture for ROCm - unit, integration, framework, workload, performance, stress, stability, scale-out, and system-level test layers - across multiple GPU generations and server platforms.
  • Define release-qualification gates and exit criteria for ROCm software releases (functional coverage, performance regressions, stability hours, scale targets, RAS criteria) and drive the org to meet them.
  • Architect the test infrastructure - distributed test runners, GitHub Actions / Jenkins / internal CI fleets, hardware lab orchestration, result data lakes, flaky-test detection, bisection automation, and self-service developer pre-submit pipelines.
  • Champion modern, agile quality engineering - shift-left testing, test pyramids, contract testing between layers, hermetic test environments, deterministic reproducers, and continuous validation in trunk.
  • Set the bar for GitHub-based quality workflows - PR gating policy, required checks, code-coverage standards, bug-bash and triage cadences, and disciplined issue management across ROCm/* repositories and partner upstream projects.
  • Lead complex escalation debug - partner with development, hardware, firmware, and customer-facing teams to root-cause the hardest multi-day, multi-node, multi-component failures and convert findings into durable test coverage.
  • Influence the roadmap - work with product management, silicon, platform, and software architecture to ensure validation readiness for next-generation Instinct GPUs and server platforms before tape-in milestones and silicon arrival.
  • Mentor and elevate Senior and Staff validation engineers, SDETs, and SQA leads; raise the technical bar through design review, code review, and written guidance.
  • Represent ROCm validation externally - strategic customer engagements, OEM qualification programs, and open-source community quality initiatives.
  • Lead system-level testing for server nodes - multi-GPU topologies, PCIe/Infinity Fabric/xGMI, BMC/IPMI, thermal/power, firmware interactions, and multi-node fabric (Ethernet/InfiniBand/UALink) bring-up and validation.Drive compute workload validation and characterization - LLM training and inference (PyTorch, vLLM, Triton, JAX), recommender systems, scientific HPC kernels, MLPerf-class benchmarks - establishing reproducible methodology, baselines, and regression tracking.


PREFERRED EXPERIENCE:

  • Software engineering experience in validation, SDET, or quality engineering, including experience leading complex systems validation.
  • Expert Python for test automation and infrastructure; strong C++ for debugging and extending production code.
  • Deep validation expertise in two or more of the following:
    • GPU software stacks (ROCm, CUDA, oneAPI, SYCL)
    • AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton, vLLM)
    • HPC runtimes and communication libraries (MPI, RCCL/NCCL, UCX, Libfabric)
    • Linux kernel, GPU drivers, or accelerator firmware
    • Distributed systems and large-scale cluster software
  • Experience validating multi-GPU, multi-node server platforms, including stress, soak, fault injection, and RAS testing.
  • Experience defining and delivering release qualification programs for hyperscalers, OEMs, or Tier-1 customers.
  • Contributions to validation, CI, or test infrastructure for ROCm, PyTorch, LLVM, Triton, vLLM, or similar open-source projects.
  • Experience leading adoption of agentic AI workflows, including automated testing, AI-driven debugging, MCP, and RAG-based engineering solutions.
  • Experience validating or operating large-scale GPU clusters (256+ GPUs), including fabric bring-up, health monitoring, and diagnostics.
  • Familiarity with AI training, inference, and HPC benchmark methodologies.
  • Experience with performance validation, profiling tools (rocprof, Omniperf, Nsight), and regression analysis.
  • Familiarity with hardware lab automation, including BMC/IPMI/Redfish, PDU control, serial consoles, automated re-imaging, and topology-aware scheduling.
  • Experience supporting validation for pre-silicon, emulation, and first-silicon accelerator bring-up.


ACADEMIC CREDENTIALS:

  • BS/MS/PhD in Computer Science, Computer Engineering, or related discipline (or equivalent demonstrated experience).


LOCATION: San Jose, California

#LI-DR1

#LI-HYBRID

About Advanced Micro Devices, Inc

Advanced Micro Devices, Inc. Careers

Join the innovative forefront of technology with a career at Advanced Micro Devices, Inc. (AMD), a leader in semiconductor development. As part of our global team, you will contribute to an organization renowned for its dedication to innovation, leadership, and diversity in the tech industry.

Work You’ll Do

At AMD, we offer job opportunities that push the boundaries of what is possible. Our team is composed of professionals who lead the way in microprocessor and graphics technology, driving industry standards and innovation. With AMD, you will be part of a culture that values growth and professional development, ensuring that every team member has the opportunity to excel.

Transform Your Career

AMD is not just about advancing technology, but also about advancing careers. Whether you are looking for an internship, a full-time position, or leadership roles, AMD provides the platform to propel your career to new heights. Our commitment to professional growth is matched by our dedication to diversity and inclusion, making AMD a place where everyone can thrive.

Innovative Work Environment

Join a team of over 12,000 dedicated professionals at the intersection of technology, industry expertise, and digital innovation. At AMD, you will work on groundbreaking projects that shape the future of computing and graphics. Our collaborative environment encourages networking and the sharing of ideas across teams and disciplines.

Career Development and Benefits

AMD is committed to the development of its employees. We offer robust training programs, including leadership development and diversity training, to ensure our team is equipped for both current challenges and future opportunities. Our benefits package is designed to support the well-being and financial security of our employees and their families.

Explore Job Opportunities

From engineering to marketing, AMD offers a range of career paths that cater to diverse skills and interests. Our hiring process is designed to be transparent and engaging, helping you to understand where you fit within our team and how you can contribute to our collective goals.

Stay Connected

Join Our Team Search open positions that match your skills and interest. We look for passionate, curious, creative, and solution-driven team players. Explore the opportunities to join a company that’s committed to your career growth and to innovation in the technology sector.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at Advanced Micro Devices, Inc.

Interview and Resume Tips

Prepare for your future with AMD by accessing resources that help you craft your resume and excel in interviews. Our goal is to help you showcase your best professional self and align your skills with the needs of our dynamic team. At Advanced Micro Devices, Inc., we empower our employees to innovate, lead, and grow. Join us in driving the future of technology while building a rewarding and sustainable career.
Learn more about Advanced Micro Devices, Inc
Size
15,500 employees
Market Cap
$100.9 billion
Industry
Net Income
$2.4 billion
Founded
1969
5 Year Trend
+30.9%
Revenue
$9.7 billion
NASDAQ

Similar Jobs

More Jobs at Advanced Micro Devices, Inc

More Information Technology Jobs

Find similar Principal Software Quality Engineer - GPU & Machine Learning jobs: