Site Reliability Engineer

xAI

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering or equivalent experience.
  • 5+ years in site reliability, systems engineering, or production operations, preferably in high-performance computing or data center environments.
  • Proven experience in incident command and technical leadership during high-pressure situations.
  • Expertise in monitoring and observability design for large-scale systems.
  • Experience across at least two domains: compute, network, storage, power, or cooling.
  • Proficient in writing and managing operational runbooks or playbooks for 24/7 environments.
  • Basic scripting skills in Python or Bash and familiarity with one systems language.

Responsibilities

  • Own the design and implementation of monitoring architecture and alert systems.
  • Provide technical leadership and coordination during severe incident responses.
  • Conduct postmortems without assigning blame, ensuring corrective actions lead to tangible improvements.
  • Lead reliability projects that cross functional boundaries amongst compute, network, and facilities.
  • Maintain high-quality runbooks and playbooks, ensuring they are up-to-date and effectively utilized during incidents.
  • Establish and monitor error budgets and availability goals aligned with business needs.
  • Engage in on-call duties and take part in responses to significant incidents at the Memphis/Southaven data center.

Benefits

  • Collaborative work environment across various engineering disciplines.
  • Opportunity to lead cross-functional reliability initiatives.
  • Engagement in high-impact incident response situations.
  • Access to cutting-edge technology in data center operations.
  • Involvement in the fast-evolving space of AI/ML infrastructure.
Full Job Description
ABOUT THE ROLE:

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries.
RESPONSIBILITIES:
  • Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure.
  • Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene.
  • Run blameless postmortems and drive corrective actions to closed, not filed.
  • Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries.
  • Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects).
  • Define error budgets and availability objectives at campus and service boundaries as adopted by the business.
  • Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus.
BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience).
  • 5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments.
  • Proven large-scale incident command experience and calm technical leadership on a bridge.
  • Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality.
  • Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry.
  • Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries.
  • Experience running game days, dependency mapping, and closed-loop corrective action programs.
  • Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry.
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

Similar Jobs

More Jobs at xAI

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: