Data Center Global Repairs Program Support

Anthropic • $320K — $405K *
US-AnywhereRemote in United States
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in data center operations with managerial or technical lead experience
  • Proven track record managing large-scale break-fix programs
  • Experience directing vendors and contract workforces toward SLAs and operational outcomes
  • Hands-on technical knowledge of server, network, and rack hardware
  • Demonstrated ability to enhance operational processes
  • Proficient with data analysis in ticketing, telemetry, and inventory management
  • Bachelor's degree in a relevant field or equivalent experience

Responsibilities

  • Define and standardize the global repair strategy and SLAs across all sites
  • Own and monitor repair turnaround times and backlog through custom dashboards
  • Develop and revise procedures for triage and return-to-service, training site partners
  • Manage RMA and reverse logistics with OEMs and ODMs, including failure analysis
  • Set parts availability to support uninterrupted repair operations across sites
  • Analyze site failure patterns to identify root causes for corrective actions
  • Conduct weekly repair reviews with vendors and site leads to ensure SLA compliance

Benefits

  • Opportunity to work in a technologically advanced, fast-paced environment
  • Collaboration with a diverse team of experts
  • Exposure to cutting-edge repair technologies especially in HPC
  • Professional development opportunities to enhance skills and knowledge
  • Flexibility with a hybrid work model requiring at least 25% on-site presence
Full Job Description


About the role

As a Repairs Lead within the Data Center Infrastructure organization, you will define and manage the end-to-end hardware repair program across Anthropic's growing fleet of data centers. You will be accountable for repair turnaround time and the compute returned to service across every site, covering server, GPU/accelerator, network, and optics break-fix, RMA and reverse logistics with OEMs and ODMs, and the spares and repair inventory that keeps repair SLAs achievable. As a subject matter expert in hardware operations, you will develop scalable repair processes and quality targets, set the standards that site operations partners and repair vendors execute against, and turn failure trends into fixes which are driven upstream with engineering and equipment partners. If you are experienced in at-scale datacenter hardware operations, are passionate about HPC data centers, & enjoy working in complex, fast-paced environments, we welcome you to apply.
Responsibilities Include:
  • Define the global repair strategy including repair SLAs, prioritization rules, escalation paths, and reporting methods, and drive standardization across all sites.
  • Own repair turnaround time and repair backlog across the fleet evidenced by Anthropic-owned ticket and telemetry dashboards you help develop.
  • Author and improve procedures for triage, break-fix, return-to-service validation, and train site operations partners on how to execute them.
  • Manage RMA and reverse logistics programs with OEMs, ODMs, and depot repair vendors, including warranty claims, return cycle times, and failure analysis feedback.
  • Set spares pool sizing and stocking levels by site and part, in coordination with supply chain and asset management, so that parts availability never gates repair SLAs.
  • Analyze failure patterns across sites to identify root causes and drive corrective actions with hardware engineering, suppliers, and site operations owners.
  • Lead the operating cadence with vendor and site leads, including weekly repair reviews, scorecards, and business reviews, and drive corrective actions for SLA excursions.
  • Communicate repair constraints, risks, and fleet availability impact to engineering and leadership.
You may be a good fit if you:
  • Have 8+ years of experience in data center operations as a manager, technical lead or related role, including accountability for production availability.
  • Can demonstrate a proven track record running break-fix programs at a large scale across multiple sites.
  • Have managed vendors, OEMs, or contract workforces to measurable outcomes: SLAs, operational reviews, and corrective action.
  • Possess hands-on technical depth in server, network, and rack-level hardware, enough to independently verify repair quality and audit vendor claims.
  • Have built or substantially improved operational processes.
  • Are comfortable working with ticket, telemetry, and inventory data to drive decisions.
  • Possess a bachelor's degree in relevant domain or equivalent practical experience.
It's a bonus if you have:
  • Experience with GPU/accelerator or high-density liquid-cooled infrastructure, including tray, cold plate, and manifold-level repair.
  • Experience managing RMA and warranty programs with hyperscale OEMs/ODMs, including failure analysis and supplier quality engagement.
  • Experience with spares planning, reverse logistics, or depot repair at data center scale.
  • Experience delivering repair outcomes inside partner-operated or colocation sites where on-the-floor operations are staffed via third parties.
  • Familiarity with optics and high-speed interconnect failures.


The annual compensation range for this role is listed below.

For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$320,000-$405,000 USD

Logistics

Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

About Anthropic

Anthropic is an artificial intelligence research lab that focuses on developing AI systems that are safe, reliable, and trustworthy. The company was founded in 2019 by Dr. Yoshua Bengio, a leading AI researcher and winner of the Turing Award. Anthropic's research is focused on developing AI systems that can learn from small amounts of data, reason about complex systems, and interact with humans in a natural way. The company is based in New York City and has a team of experienced AI researchers and engineers.
Learn more about Anthropic
Size
50 employees
Industry
Founded
2019

Similar Jobs

More Jobs at Anthropic

More Technical Services Jobs

Find similar Data Center Global Repairs Program Support jobs: