ABOUT THE ROLE:As the Director of Site Operations, you'll own node and rack uptime for SpaceXAI's AI supercompute cluster-the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You'll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We're looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow.
RESPONSIBILITIES:- Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster.
- Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org.
- Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum.
- Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime.
- Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires.
- Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation.
- Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites.
- Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners.
- Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales.
BASIC QUALIFICATIONS:- Bachelor's degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams OR 10+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams.
PREFERRED SKILLS AND EXPERIENCE:- Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings.
- Deep expertise in server hardware, cluster reliability, and data center technologies, from deployment through lifecycle management.
- Experience supporting compute-heavy environments like AI, machine learning, or high-performance computing at scale.
- A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute clusters.
- Experience leading site reliability engineering or equivalent reliability-focused teams, including root cause analysis and procedure ownership.
- Strong analytical skills and the ability to explain technical concepts clearly to diverse audiences, from technicians to executive and tenant partners.
- A history of partnering with vendors at scale, driving mean time to repair down, and scaling operations across multiple sites.
- Familiarity with tooling and automation (e.g., Jira, Python, Bash) used to monitor cluster health and improve team efficiency.
- Enthusiasm for SpaceXAI's mission to accelerate human discovery and unravel the universe.
- Ability to thrive in a dynamic, mission-focused environment with on-call ownership of cluster-impacting events.
ADDITIONAL REQUIREMENTS:- Willingness to travel frequently to data center locations to support operations across sites.
- Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use.
- Must be willing to work extended hours and/or weekends as needed.