Lead, Hardware Deployment Engineer

SpaceXAI

$120K — $145K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience deploying or repairing data center hardware
  • Experience in L11 integration for GPU systems
  • Proven leadership in fast-paced deployment or manufacturing environments
  • Strong troubleshooting skills across various hardware systems
  • Willing to work on-site in Memphis, TN, during critical phases.

Responsibilities

  • Lead and build a hardware deployment team including engineers and technicians
  • Integrate and bring up GPU hardware across multiple data halls concurrently
  • Achieve high node availability shortly after hardware delivery
  • Maintain post-integration hardware health and availability
  • Internalize repairs to minimize turnaround times and backlog
  • Define and enforce vendor SLAs to prevent repair backlogs
  • Analyze hardware failures and coordinate corrective actions with vendors
  • Collaborate with site operations to streamline hardware deployment processes.

Benefits

  • Opportunity to lead a critical deployment team in a cutting-edge technology environment
  • Direct impact on the operational efficiency of AI compute clusters
  • Chance to work with advanced GPU technologies and data center systems
  • Involvement in process improvement and innovation in hardware deployment
  • Potential for career growth and development in a rapidly evolving field.
Full Job Description
ABOUT THE ROLE:

As the Hardware Deployment Engineer Lead, you will own the end-to-end bring-up of GPU compute hardware across the world's largest AI training clusters. You will build and lead a dedicated in-house hardware deployment team responsible for L11 integration, hardware bring-up, and post-L11 repair of GB300-class systems across multiple data halls concurrently. Your team's throughput directly determines how fast xAI can compute online - this is one of the most critical path activities in the company. You will set the deployment playbook, hold hardware vendors accountable to SLAs, and institutionalize processes so that cluster deployment is limited only by hardware supply and power, never by deployment velocity. The position is based in Memphis, TN.
RESPONSIBILITIES:
  • Lead, hire, and develop a dedicated hardware deployment team (deployment engineers, deployment technicians, and repair technicians) with full ownership of team structure and staffing.
  • Own L11 rack integration and compute hardware bring-up across multiple data halls concurrently, from delivery dock to healthy production handoff.
  • Drive aggressive bring-up timelines: achieve 95%+ node availability within days of rack delivery and 100% closure within one week per data hall.
  • Own post-L11 hardware health: run systematic health pushes to sustain greater than 98% node availability prior to turnover to operations.
  • Internalize non-RMA hardware repairs to maximize hardware recovery, minimize repair backlogs, and reduce dependence on OEM turnaround times.
  • Develop and enforce vendor SLAs for OEM and supplier responsibilities; prevent accumulation of unrepaired hardware ("bone piles") and repair backlogs before turnover to operations.
  • Perform root cause analysis of hardware failures discovered during L11 and drive corrective actions with vendors and internal engineering teams.
  • Partner with site operations on hardware debugging and repair, and train site operations teams to support future data center deployments.
  • Build, document, and continuously improve deployment processes, tooling, and training so bring-up capability scales across sites and future hardware generations.
BASIC QUALIFICATIONS:
  • 5+ years of hands-on experience deploying, integrating, or repairing compute/server hardware at data center scale.
  • Direct experience with L11 (rack-level) integration and bring-up of GPU or accelerator-based systems.
  • Demonstrated experience leading technician or engineering teams in a fast-paced deployment, manufacturing, or data center environment.
  • Deep troubleshooting skills across servers, GPUs, NVLink/fabric interconnects, high-speed networking, and liquid cooling systems.
  • Willingness to work on-site in Memphis, TN, including extended hours and weekends during critical bring-up phases.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience with NVIDIA GB200/GB300 NVL72 or similar rack-scale liquid-cooled GPU systems.
  • Experience standing up a new team or function, including hiring, training, and process development from scratch.
  • Experience managing OEM/ODM vendor relationships (e.g., Dell, Supermicro), including SLA definition and enforcement.
  • Experience with hardware failure analysis, RMA processes, and component-level repair strategies at fleet scale.
  • Experience with data center automation, burn-in/validation tooling, and hardware health telemetry.
  • Track record of driving step-change improvements in deployment velocity or cost (e.g., insourcing work previously performed by OEMs).

Similar Jobs

More Jobs at SpaceXAI

More Technical Services Jobs

Find similar Lead, Hardware Deployment Engineer jobs: