Advanced Micro Devices, Inc

Principal Cluster Reliability Architect

Advanced Micro Devices, Inc$160K — $190K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-10 years of experience in systems architecture, particularly in distributed systems and reliability engineering.
  • Proven expertise in large-scale infrastructure, AI, and HPC environments.
  • Strong background in Linux systems and production-scale platforms.
  • Experience with Kubernetes, Slurm, and resiliency frameworks like chaos engineering.
  • Ability to influence teams and drive architectural change without direct authority.

Responsibilities

  • Define a comprehensive cluster reliability framework.
  • Establish reliability requirements for architectures and deployment models.
  • Drive architectural enhancements for fault tolerance and system recoverability.
  • Collaborate with SRE and Platform Operations teams to model sustainable operations.
  • Create standards for observability and health monitoring.
  • Implement validation strategies for reliability at production scale.
  • Ensure consistent reliability practices across the entire cluster lifecycle.

Benefits

  • Comprehensive AMD benefits package including health and wellness programs.
  • Opportunities for professional development and career advancement.
  • Flexible work arrangements to promote work-life balance.
  • Participation in a collaborative and innovative corporate culture.
Full Job Description
THE ROLE

AMD is seeking a Principal Cluster Reliability Architect to define and drive the reliability strategy for next-generation AI and HPC cluster platforms. This role serves as the technical authority for end-to-end cluster reliability, spanning compute, networking, storage, and control plane architectures. The successful candidate will influence architecture, deployment, validation, and operational readiness across AMD reference designs and large-scale customer deployments.

The Principal Cluster Reliability Architect will establish reliability frameworks, drive cross-functional alignment, and ensure that reliability considerations are embedded throughout the cluster lifecycle from architecture through production operations.

THE PERSON

You are a highly experienced systems architect with deep expertise in distributed systems, large-scale infrastructure, and reliability engineering. You possess a strong understanding of AI/HPC environments and have demonstrated success designing and operating production-scale platforms.

You excel at leading through influence, aligning diverse engineering teams around common objectives, and translating operational learnings into architectural improvements. You combine technical depth with strategic thinking and are passionate about delivering resilient, scalable, and operationally efficient infrastructure platforms.

KEY RESPONSIBILITIES

Reliability Architecture & Design
  • Define and implement a comprehensive cluster reliability architecture framework spanning the full infrastructure lifecycle.
  • Establish reliability requirements for AMD reference architectures and customer deployment models.
  • Drive architecture decisions that enhance fault tolerance, redundancy, failure isolation, and system recoverability.
  • Influence hardware, firmware, software, and operational design decisions to improve system resiliency.

Operational Readiness & Day-2 Operations
  • Partner with Site Reliability Engineering (SRE) and Platform Operations teams to define sustainable operational models.
  • Establish standards for:
    • Observability and telemetry
    • Health monitoring frameworks
    • Job-aware failure detection and recovery mechanisms
    • Lifecycle management, including patching and upgrades
  • Create closed-loop feedback mechanisms that convert operational insights into architectural improvements.

Validation & Reliability Engineering
  • Define reliability KPIs, SLAs, SLOs, and deployment acceptance criteria.
  • Develop and execute reliability validation strategies at production scale.
  • Establish methodologies for:
    • Failure injection testing
    • Chaos engineering
    • Recovery validation
    • Resiliency benchmarking
  • Drive data-driven reliability improvements through empirical testing and analysis.

Lifecycle Integration & Process Maturity
  • Ensure reliability consistency across:
    • Architecture
    • Deployment
    • Cluster bring-up
    • Validation
    • Operations
  • Identify and eliminate gaps in ownership, governance, handoff processes, and operational readiness.
  • Drive repeatable engineering practices that improve scalability and deployment quality.

Cross-Functional Technical Leadership
  • Act as AMD's subject matter expert for cluster reliability architecture.
  • Lead alignment across Architecture, Deployment Engineering, Validation Engineering, Platform Operations, and Product organizations.
  • Mentor senior engineers and architects.
  • Represent reliability strategy in customer engagements and strategic programs.
  • Influence product and platform roadmaps to improve reliability outcomes across AMD cluster initiatives.

PREFERRED SKILLS

Technical Expertise
  • Large-scale cluster architecture across compute, networking, storage, and control plane systems.
  • Distributed systems architecture, reliability engineering, and failure modeling.
  • Linux infrastructure and systems software.
  • Production-scale cloud, AI, HPC, or platform infrastructure.

AI/HPC Technologies
  • Kubernetes
  • Slurm
  • ROCm
  • GPU cluster architectures
  • AI training and inference infrastructure

Networking & Platform Technologies
  • RDMA networking technologies:
    • InfiniBand
    • RoCE
  • High-performance cluster interconnects
  • Distributed storage architectures

Reliability Engineering
  • Site Reliability Engineering (SRE) methodologies
  • Observability platforms and telemetry systems
  • Incident management and root cause analysis
  • Automation and infrastructure-as-code practices
  • Chaos engineering and resilience testing

Leadership & Influence
  • Development of reference architectures and platform standards
  • Cross-functional technical leadership
  • Executive-level communication and stakeholder management
  • Experience influencing technical direction across large organizations without direct authority


ACADEMIC CREDENTIALS
  • Bachelor's degree in Computer Engineering, Computer Science, Electrical Engineering, or a related technical discipline.


This role is not eligible for visa sponsorship.

#LI-KW1

Benefits offered are described: AMD benefits at a glance.

About Advanced Micro Devices, Inc

Advanced Micro Devices, Inc. Careers

Join the innovative forefront of technology with a career at Advanced Micro Devices, Inc. (AMD), a leader in semiconductor development. As part of our global team, you will contribute to an organization renowned for its dedication to innovation, leadership, and diversity in the tech industry.

Work You’ll Do

At AMD, we offer job opportunities that push the boundaries of what is possible. Our team is composed of professionals who lead the way in microprocessor and graphics technology, driving industry standards and innovation. With AMD, you will be part of a culture that values growth and professional development, ensuring that every team member has the opportunity to excel.

Transform Your Career

AMD is not just about advancing technology, but also about advancing careers. Whether you are looking for an internship, a full-time position, or leadership roles, AMD provides the platform to propel your career to new heights. Our commitment to professional growth is matched by our dedication to diversity and inclusion, making AMD a place where everyone can thrive.

Innovative Work Environment

Join a team of over 12,000 dedicated professionals at the intersection of technology, industry expertise, and digital innovation. At AMD, you will work on groundbreaking projects that shape the future of computing and graphics. Our collaborative environment encourages networking and the sharing of ideas across teams and disciplines.

Career Development and Benefits

AMD is committed to the development of its employees. We offer robust training programs, including leadership development and diversity training, to ensure our team is equipped for both current challenges and future opportunities. Our benefits package is designed to support the well-being and financial security of our employees and their families.

Explore Job Opportunities

From engineering to marketing, AMD offers a range of career paths that cater to diverse skills and interests. Our hiring process is designed to be transparent and engaging, helping you to understand where you fit within our team and how you can contribute to our collective goals.

Stay Connected

Join Our Team Search open positions that match your skills and interest. We look for passionate, curious, creative, and solution-driven team players. Explore the opportunities to join a company that’s committed to your career growth and to innovation in the technology sector.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at Advanced Micro Devices, Inc.

Interview and Resume Tips

Prepare for your future with AMD by accessing resources that help you craft your resume and excel in interviews. Our goal is to help you showcase your best professional self and align your skills with the needs of our dynamic team. At Advanced Micro Devices, Inc., we empower our employees to innovate, lead, and grow. Join us in driving the future of technology while building a rewarding and sustainable career.
Learn more about Advanced Micro Devices, Inc
Size
15,500 employees
Market Cap
$100.9 billion
Industry
Net Income
$2.4 billion
Founded
1969
5 Year Trend
+30.9%
Revenue
$9.7 billion
NASDAQ

Similar Jobs

More Jobs at Advanced Micro Devices, Inc

More Information Technology Jobs

Find similar Principal Cluster Reliability Architect jobs: