Advanced Micro Devices, Inc

Cloud & Customer Solutions Engineer - DC GPU

Advanced Micro Devices, Inc$130K — $155K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in production software or infrastructure engineering
  • Hands-on GPU compute experience at scale
  • Strong knowledge of AI infrastructure stack and observability tools
  • Familiarity with cloud platforms like AWS, Azure, or GCP
  • Proficient in Python and one systems language
  • Direct experience with customer-facing technical support
  • Experience with ROCm and AMD Instinct GPUs is a plus

Responsibilities

  • Own customer deployments from start to finish
  • Deploy and tune large-scale training and inference stacks
  • Lead incident analysis and resolution for production issues
  • Implement AI solutions and maintain their operational behavior
  • Build tools for observability and production readiness certification
  • Transfer operational knowledge to customer teams
  • Contribute field insights to enhance internal knowledge bases
  • Convert findings into contributions to AMD's product and software teams

Benefits

  • Comprehensive AMD benefits package
  • Flexible hybrid work environment
  • Opportunities for professional development
  • Access to cutting-edge technology and tools
  • Collaborative team culture with industry leaders
Full Job Description
THE TEAM:

AMD's Data Center GPU organization is transforming the industry with our AI based Graphic Processors. Our primary objective is to design exceptional products that drive the evolution of computing experiences, serving as the cornerstone for enterprise Data Centers, (AI) Artificial Intelligence, HPC and Embedded systems. If this resonates with you, come and joining our Data Center GPU organization where we are building amazing AI powered products with amazing people.

THE ROLE:

As a Cloud and Customer Solutions Engineer on AMD's Applied AI team, you will embed directly with AMD's most strategic AI customers - frontier labs, NeoCloud providers, CSPs, and AI-native companies - to take AMD Instinct GPU clusters from delivery to sustained production excellence. You own the customer outcome end-to-end: cluster bring-up and certification, workload deployment and performance, production incident response, and the transfer of operational capability that moves customers toward autonomous operation of their AMD fleets.

To be direct about what this role is: despite the "Solutions" title, this is not a pre-sales or demo role. You will write production code, operate live clusters, carry accountability for customer production outcomes, and be the engineer in the room when things break at scale. What you learn in the field, you convert into durable improvements - to ROCm, to the open-source serving ecosystem, and to the reference architectures every subsequent deployment inherits.

THE PERSON:

You are a strong production engineer who is energized rather than drained by ambiguity, customer pressure, and environments you do not control. You can debug a distributed training hang at 2am, explain the root cause to a customer VP at 9am, and land the fix upstream by the end of the week. You measure success by customer production outcomes, not code merged or tickets closed. When something is broken on a cluster you touch, it is your problem until it is fixed or explicitly handed off.

KEY RESPONSIBILITIES:
  • Own customer deployments end-to-end: cluster bring-up and burn-in, production readiness certification, workload onboarding, performance validation, and sustained production operation on AMD Instinct GPU fleets
  • Deploy and tune large-scale training and inference stacks (ROCm, vLLM, SGLang, RCCL, Kubernetes, Slurm) against customer-specific workloads and SLOs across cloud, NeoCloud, and bare-metal environments
  • Lead root-cause analysis and resolution of production incidents on customer clusters, including Sev-1 response, and drive fixes to permanent closure
  • Deploy agentic AI solutions into customer environments in partnership with Agentic Data Engineers, and own their production behavior within the engagement
  • Build the observability, benchmarking, and validation tooling needed to certify clusters as production-ready and keep them there
  • Transfer operational capability to customer teams - documentation, runbooks, and hands-on enablement - moving customers up the operator-autonomy ladder from assisted operation to independent production ownership
  • Contribute field learnings to the Applied AI team's skills library and engagement memory databases, so deployment knowledge compounds across the practice
  • Convert field findings into upstream contributions - ROCm issues and patches, serving-framework improvements, reference-architecture updates - and provide structured field signal to AMD product, software, and silicon teams

PREFERRED EXPERIENCE:
  • 5+ years of production software or infrastructure engineering, including significant time operating or deploying systems in environments you did not build (level flexible for exceptional candidates)
  • Hands-on experience with GPU compute at scale: cluster deployment, distributed training or high-throughput inference, performance debugging, and workload optimization
  • Strong working knowledge of the modern AI infrastructure stack: Kubernetes and/or Slurm, containerized GPU workloads, collective communication libraries (RCCL/NCCL), high-performance networking (RoCE/InfiniBand), and observability tooling (Prometheus, Grafana)
  • Cloud platform depth (AWS, Azure, GCP, or NeoCloud environments), including hybrid and bare-metal deployment patterns
  • Proficiency in Python and at least one systems language; comfort navigating and modifying large codebases you did not write
  • Working familiarity with LLM application patterns - inference serving, RAG, and agentic workflows - sufficient to deploy and troubleshoot them in customer environments
  • Direct customer-facing experience: embedded deployments, technical escalations, on-site engagements, or equivalent
  • Experience with ROCm and AMD Instinct GPUs strongly preferred; deep CUDA-ecosystem experience with demonstrated ability to work cross-platform also valued
  • Open-source contribution history in AI/ML infrastructure projects is a plus

PREFERRED ACADEMIC CREDENTIALS:
  • Bachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent practical experience

This role is not eligible for visa sponsorship.

#LI-RW1

#LI-HYBRID

Benefits offered are described: AMD benefits at a glance.

About Advanced Micro Devices, Inc

Advanced Micro Devices, Inc. Careers

Join the innovative forefront of technology with a career at Advanced Micro Devices, Inc. (AMD), a leader in semiconductor development. As part of our global team, you will contribute to an organization renowned for its dedication to innovation, leadership, and diversity in the tech industry.

Work You’ll Do

At AMD, we offer job opportunities that push the boundaries of what is possible. Our team is composed of professionals who lead the way in microprocessor and graphics technology, driving industry standards and innovation. With AMD, you will be part of a culture that values growth and professional development, ensuring that every team member has the opportunity to excel.

Transform Your Career

AMD is not just about advancing technology, but also about advancing careers. Whether you are looking for an internship, a full-time position, or leadership roles, AMD provides the platform to propel your career to new heights. Our commitment to professional growth is matched by our dedication to diversity and inclusion, making AMD a place where everyone can thrive.

Innovative Work Environment

Join a team of over 12,000 dedicated professionals at the intersection of technology, industry expertise, and digital innovation. At AMD, you will work on groundbreaking projects that shape the future of computing and graphics. Our collaborative environment encourages networking and the sharing of ideas across teams and disciplines.

Career Development and Benefits

AMD is committed to the development of its employees. We offer robust training programs, including leadership development and diversity training, to ensure our team is equipped for both current challenges and future opportunities. Our benefits package is designed to support the well-being and financial security of our employees and their families.

Explore Job Opportunities

From engineering to marketing, AMD offers a range of career paths that cater to diverse skills and interests. Our hiring process is designed to be transparent and engaging, helping you to understand where you fit within our team and how you can contribute to our collective goals.

Stay Connected

Join Our Team Search open positions that match your skills and interest. We look for passionate, curious, creative, and solution-driven team players. Explore the opportunities to join a company that’s committed to your career growth and to innovation in the technology sector.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Personalize your subscription to receive job alerts, latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at Advanced Micro Devices, Inc.

Interview and Resume Tips

Prepare for your future with AMD by accessing resources that help you craft your resume and excel in interviews. Our goal is to help you showcase your best professional self and align your skills with the needs of our dynamic team. At Advanced Micro Devices, Inc., we empower our employees to innovate, lead, and grow. Join us in driving the future of technology while building a rewarding and sustainable career.
Learn more about Advanced Micro Devices, Inc
Size
15,500 employees
Market Cap
$100.9 billion
Industry
Net Income
$2.4 billion
Founded
1969
5 Year Trend
+30.9%
Revenue
$9.7 billion
NASDAQ

Similar Jobs

More Jobs at Advanced Micro Devices, Inc

More Enterprise Technology Jobs

Find similar Cloud & Customer Solutions Engineer - DC GPU jobs: