Director, Site Operations

xAI

$150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree with 7+ years in large scale operations and 5+ years leading technical teams, or 10+ years in large scale operations with 5+ years leading people leaders.
  • Proven experience in 24/7 operations in high-pressure environments.
  • Expertise in server hardware and data center technology lifecycle management.
  • Background in supporting high-performance compute environments like AI or machine learning.
  • Strong analytical skills for data-driven decision making and incident management.
  • Experience collaborating with vendors to optimize repair and operational efficiency.
  • Familiarity with automation tools such as Jira, Python, or Bash.

Responsibilities

  • Own and ensure exceptional uptime for SpaceXAI's supercompute cluster across multiple sites.
  • Lead and foster a high-performance culture within a 250+ person operations team.
  • Drive swift remediation of node and rack failures through both technical and physical intervention.
  • Collaborate with various departments to mitigate downtime through effective maintenance and upgrades.
  • Oversee vendor operations for hardware rework and repairs efficiently.
  • Lead site reliability engineering efforts to monitor cluster health and mitigate faults effectively.
  • Champion continuous improvement initiatives using operational data to enhance SLA performance.

Benefits

  • Healthcare coverage including medical, dental, and vision.
  • 401(k) plan with matching contributions.
  • Generous paid time off policy and holidays.
  • Professional development opportunities and training programs.
  • Camaraderie in a mission-driven work environment with a focus on innovation.
Full Job Description
ABOUT THE ROLE:

As the Director of Site Operations, you'll own node and rack uptime for SpaceXAI's AI supercompute cluster-the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You'll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We're looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow.

RESPONSIBILITIES:
  • Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster.
  • Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org.
  • Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum.
  • Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime.
  • Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires.
  • Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation.
  • Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites.
  • Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners.
  • Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales.

BASIC QUALIFICATIONS:
  • Bachelor's degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams OR 10+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams.

PREFERRED SKILLS AND EXPERIENCE:
  • Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings.
  • Deep expertise in server hardware, cluster reliability, and data center technologies, from deployment through lifecycle management.
  • Experience supporting compute-heavy environments like AI, machine learning, or high-performance computing at scale.
  • A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute clusters.
  • Experience leading site reliability engineering or equivalent reliability-focused teams, including root cause analysis and procedure ownership.
  • Strong analytical skills and the ability to explain technical concepts clearly to diverse audiences, from technicians to executive and tenant partners.
  • A history of partnering with vendors at scale, driving mean time to repair down, and scaling operations across multiple sites.
  • Familiarity with tooling and automation (e.g., Jira, Python, Bash) used to monitor cluster health and improve team efficiency.
  • Enthusiasm for SpaceXAI's mission to accelerate human discovery and unravel the universe.
  • Ability to thrive in a dynamic, mission-focused environment with on-call ownership of cluster-impacting events.

ADDITIONAL REQUIREMENTS:
  • Willingness to travel frequently to data center locations to support operations across sites.
  • Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use.
  • Must be willing to work extended hours and/or weekends as needed.

Similar Jobs

More Jobs at xAI

More Enterprise Technology Jobs

Find similar Director, Site Operations jobs: