Data Center Operations Lead - Partner Site Operations

Anthropic$320K — $405K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in data center operations, including production availability accountability.
  • Proven experience managing vendors and contract workforces to achieve measurable outcomes.
  • Hands-on technical depth in server, network, and rack-level infrastructure for quality audits.
  • Experience in building or significantly improving operational processes.
  • Previous role in incident command or as lead responder, able to communicate effectively under pressure.
  • Willingness to support non-standard hours and participate in on-call rotations.
  • Bachelor's degree in a relevant field or equivalent practical experience.

Responsibilities

  • Own operational outcomes including site availability and deployment milestones.
  • Set daily and weekly vendor priorities and lead operational cadence activities.
  • Define and update processes for deployment and compliance across sites.
  • Monitor vendor performance against SLAs and initiate corrective actions when needed.
  • Manage incident response during escalations and direct vendor communications.
  • Translate engineering needs into actionable vendor directives and communicate risk to management.

Benefits

  • Opportunities for professional development and continuous learning.
  • Visa sponsorship available for qualified candidates.
  • Participation in operational readiness initiatives for new infrastructure.
Full Job Description
About the role

Anthropic's Data Center Operations (DCO) team ensures compute fleet availability through hardware and IT operations. At our partner-operated sites, this role manages the interface between Anthropic and the strategic site operations partner performing day-to-day data hall work.

As the site lead, you own site outcomes for your assigned sites including: deployment velocity, availability, and incident response. Rather than managing operations staff directly, you provide tactical direction, set priorities, and define the standards for the vendor's on-site teams, paired with performance oversight and ongoing operational assessment to ensure all operational commitments are met.

You will define the operational processes, quality gates, and governance rhythms for partner-operated sites. Expect to build the playbook as much as you run it, not just at a site level, but defining and developing program improvements fleet-wide.

What you'll own
  • Operational outcomes. Own site availability, deployment milestones, and repair turnaround, verified with independent data rather than vendor self-reporting.
  • Vendor direction. Set daily and weekly priorities and lead the operating cadence, including standups and business reviews.
  • Process definition. Author and improve procedures for deployment, break-fix, change management, security, and EHS compliance. Analyze operational trends and standardize lessons across the program.
  • Performance management. Track vendor performance against SLAs and staffing commitments, driving corrective actions when necessary.
  • Incident response and on-call. Participate in the incident escalation on-call rotation. When designated Anthropic Incident Commander for a site-specific incident, direct vendor response, own communications, and close out post-incident actions.
  • Internal interface. Translate engineering requirements into vendor direction and communicate site constraints and risks to leadership.


Representative work
  • Leading weekly operations reviews and scorecards with vendor site leads.
  • Directing deployment surges to meet first-compute-online milestones.
  • Analyzing failure patterns to identify root causes and driving fixes with owners.
  • Creating break-fix ownership matrices and training vendor teams.
  • Serving as Incident Commander for facility events and producing post-mortems.
  • Establishing operational readiness for new data halls, including spares and security.
  • Identifying process gaps and codifying improvements as program standards.


You may be a good fit if you
  • Have 8+ years of experience in data center operations (hardware, IT infrastructure, or critical facilities) as a manager, technical lead or related role, including accountability for production availability.
  • Have managed vendors, MSPs, or contract workforces to measurable outcomes: SOWs, SLAs, operational reviews, and corrective action.
  • Carry hands-on technical depth in server, network, and rack-level infrastructure, enough to independently verify vendor claims and audit quality.
  • Have built or substantially improved operational processes, not just run them.
  • Have served in an incident command or lead-responder role and communicate clearly under ambiguity.
  • Can support non-standard hours, including an on-call rotation and availability during deployment surges and maintenance windows.
  • Bachelor's degree in relevant domain or equivalent practical experience.
Strong candidates may also have
  • Experience with third-party colocation providers or partner-operated sites, delivering IT operations outcomes inside a facility someone else runs.
  • Experience standing up operations at a new site or data hall, from commissioning handoff through first deployment.
  • Experience with GPU/accelerator or high-density liquid-cooled infrastructure.
  • Familiarity with multi-vendor sites where facilities and IT operations are performed by different partners.
  • Experience leading projects from initiation to completion across teams you didn't own.
  • Background in incident management frameworks, contract/SLA design, or EHS programs.


The annual compensation range for this role is listed below.

For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.

Annual Salary:

$320,000-$405,000 USD

Logistics

Minimum education: Bachelor's degree or an equivalent combination of education, training, and/or experience

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position

Location-based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

About Anthropic

Anthropic is an artificial intelligence research lab that focuses on developing AI systems that are safe, reliable, and trustworthy. The company was founded in 2019 by Dr. Yoshua Bengio, a leading AI researcher and winner of the Turing Award. Anthropic's research is focused on developing AI systems that can learn from small amounts of data, reason about complex systems, and interact with humans in a natural way. The company is based in New York City and has a team of experienced AI researchers and engineers.
Learn more about Anthropic
Size
50 employees
Industry
Founded
2019

Similar Jobs

More Jobs at Anthropic

More Information Technology Jobs

Find similar Data Center Operations Lead - Partner Site Operations jobs: