Bachelor's degree with 7+ years in large scale operations and 5+ years leading technical teams, or 10+ years in large scale operations with 5+ years leading people leaders.
Proven experience in 24/7 operations in high-pressure environments.
Expertise in server hardware and data center technology lifecycle management.
Background in supporting high-performance compute environments like AI or machine learning.
Strong analytical skills for data-driven decision making and incident management.
Experience collaborating with vendors to optimize repair and operational efficiency.
Familiarity with automation tools such as Jira, Python, or Bash.
Responsibilities
Own and ensure exceptional uptime for SpaceXAI's supercompute cluster across multiple sites.
Lead and foster a high-performance culture within a 250+ person operations team.
Drive swift remediation of node and rack failures through both technical and physical intervention.
Collaborate with various departments to mitigate downtime through effective maintenance and upgrades.
Oversee vendor operations for hardware rework and repairs efficiently.
Lead site reliability engineering efforts to monitor cluster health and mitigate faults effectively.
Champion continuous improvement initiatives using operational data to enhance SLA performance.
Benefits
Healthcare coverage including medical, dental, and vision.
401(k) plan with matching contributions.
Generous paid time off policy and holidays.
Professional development opportunities and training programs.
Camaraderie in a mission-driven work environment with a focus on innovation.
Full Job Description
ABOUT THE ROLE:
As the Director of Site Operations, you'll own node and rack uptime for SpaceXAI's AI supercompute cluster-the most advanced of its kind. This role is the extreme owner of cluster health and customer Service Level Agreements across 5+ sites operating 24/7. You'll lead a 250+ person organization of site managers, shift supervisors, and technicians, plus the site reliability engineering team that monitors cluster health and drives fault mitigation at scale. We're looking for a hands-on operations leader who can build a culture of excellence and accountability, partner tightly across the company, and keep uptime exceptional as we grow.
RESPONSIBILITIES:
Own Cluster Uptime: Serve as extreme owner of node, rack, and cluster health across 5+ sites running 24/7, accountable for customer Service Level Agreements and consistently exceptional uptime on SpaceXAI's supercompute cluster.
Lead a Large Operations Organization: Direct a 250+ person team spanning site managers, shift supervisors, and technicians across four 24/7 shifts, building a culture of excellence and accountability at every layer of the org.
Drive Node and Rack Remediation: Ensure systematic recovery of failed nodes and racks through command-line and physical intervention, driving mean time to repair to the feasible minimum.
Partner Across Functions: Coordinate with facilities operations to limit downtime from power and cooling faults and proactive maintenance; with network engineering on cluster upgrades; and with tenant representatives on node remediation and planned and unplanned downtime.
Own Vendor Execution: Direct vendors through hardware rework and field operations so repairs, replacements, and capacity work happen at the speed the cluster requires.
Lead Site Reliability Engineering: Own the SRE organization responsible for proactive cluster health monitoring, reactive fault mitigation at scale, root cause analyses for node, rack, and cluster issues, and site-wide reliability procedures and fault documentation.
Run Data-Driven Improvement: Lead continual improvement and efficiency initiatives, using operational data to balance team resources and raise uptime, repair time, and SLA performance across sites.
Command Incidents at Scale: Set the standard for incident response during cluster-impacting events, providing clear direction, fast recovery, and tight communication with internal and external partners.
Scale Operations: Standardize best practices across sites and grow the organization in step with cluster expansion, keeping operations consistent as SpaceXAI's footprint scales.
BASIC QUALIFICATIONS:
Bachelor's degree and 7+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams OR 10+ years of experience working in a large scale operations with 5+ years leading people leaders of technical teams.
PREFERRED SKILLS AND EXPERIENCE:
Proven ability to lead large, multi-site, 24/7 operations organizations in fast-paced, high-responsibility settings.
Deep expertise in server hardware, cluster reliability, and data center technologies, from deployment through lifecycle management.
Experience supporting compute-heavy environments like AI, machine learning, or high-performance computing at scale.
A track record of owning uptime, Service Level Agreements, or reliability metrics for large compute clusters.
Experience leading site reliability engineering or equivalent reliability-focused teams, including root cause analysis and procedure ownership.
Strong analytical skills and the ability to explain technical concepts clearly to diverse audiences, from technicians to executive and tenant partners.
A history of partnering with vendors at scale, driving mean time to repair down, and scaling operations across multiple sites.
Familiarity with tooling and automation (e.g., Jira, Python, Bash) used to monitor cluster health and improve team efficiency.
Enthusiasm for SpaceXAI's mission to accelerate human discovery and unravel the universe.
Ability to thrive in a dynamic, mission-focused environment with on-call ownership of cluster-impacting events.
ADDITIONAL REQUIREMENTS:
Willingness to travel frequently to data center locations to support operations across sites.
Physical capability to handle data center tasks, including lifting up to 50 lbs unassisted, standing for long periods, and occasional ladder use.
Must be willing to work extended hours and/or weekends as needed.