THE ROLEAMD is seeking a highly technical leader to support platform architecture, serviceability, telemetry, and associated debug across AMD Instinct™ AI platforms. This role combines architecture ownership with hands-on execution, root-cause analysis, and deployment readiness across BMC, platform firmware, system software, and management infrastructure.
THE PERSONOwn the architecture, execution, and technical leadership of AMD Instinct platform management and serviceability solutions, serving as the subject matter expert for BMC, Redfish, telemetry, diagnostics, observability, and platform-level debug while ensuring deployment readiness and customer success.
Success is measured by architecture delivery, faster root-cause resolution, improved serviceability, enhanced telemetry coverage, and reduced customer-impacting issues.
KEY RESPONSIBILITIES- Own architecture for platform management, serviceability, diagnostics, and observability.
- Define and drive BMC, Redfish, telemetry, inventory management, and health monitoring solutions.
- Lead cross-functional debug and root-cause analysis of platform issues spanning hardware, firmware, BMC, and software.
- Drive customer escalations, deployment readiness, and field issue resolution.
- Establish diagnostic frameworks, debug methodologies, and serviceability requirements.
- Partner with silicon, platform, firmware, software, validation, and customer engineering teams.
- Influence future platform management architecture and industry standards adoption.
REQUIRED EXPERIENCE- Strong background in server, AI accelerator, GPU, or datacenter platforms.
- Deep expertise in BMC architecture and OpenBMC environments.
- Hands-on experience with Redfish, PLDM, MCTP, IPMI, and DMTF standards.
- Knowledge of I2C, SMBus, PMBus, SPI, PCIe, UART, and platform sideband interfaces.
- Experience in telemetry, sensors, event logging, diagnostics, and root-cause analysis.
- Proven ability to lead complex technical investigations and drive issue closure across organizations.
PREFERRED EXPERIENCE- Experience with hyperscale datacenter deployments.
- Understanding of RAS, fault management, and serviceability architectures.
- Strong automation and debug tooling experience.
- Experience supporting customer escalations and production platforms.
ACADEMIC CREDENTIALSBachelor's or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field. Advanced degree preferred.
This role is not eligible for visa sponsorship.#LI-BW2
#LI-HYBRID
Benefits offered are described: AMD benefits at a glance.