Sr Staff Site Reliability Engineer, AI Infrastructure

d-Matrix

• $150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master’s in Computer Science, Electrical Engineering, or related field; 7+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux systems knowledge with hands-on experience in colocation or on-premises server infrastructure.
  • Production IaC experience with Terraform and/or Ansible, focusing on writing and maintaining configurations.
  • Kubernetes operational experience for cluster troubleshooting and workload management.
  • Familiarity with observability tools (e.g., Prometheus/Grafana, DataDog, Splunk) for building dashboards and writing alert rules.
  • Proficient in Python and/or Bash scripting, with incident response experience.

Responsibilities

  • Own reliability and availability across colo server fleets, on-premises lab clusters, and cloud environments.
  • Perform hands-on infrastructure work including server provisioning, OS configuration, and troubleshooting.
  • Lead capacity planning and manage hardware lifecycle while tracking cloud expenditures.
  • Drive provisioning and operational changes using Terraform and/or Ansible, contributing to shared modules.
  • Build automation to eliminate manual tasks, including fleet health checks and networking automation.
  • Design monitoring and alerting systems, contributing to AIOps detection workflows.
  • Participate in on-call rotations and produce high-quality root cause analyses for incidents.
  • Support platform services for internal teams and customers, ensuring quality and uptime.

Benefits

  • Collaborative work environment with opportunities for professional growth.
  • Hands-on project ownership leading impactful infrastructure improvements.
  • Exposure to a wide range of cutting-edge technologies in AI and cloud infrastructures.
Full Job Description
Role Overview

d-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on - colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. This is hands-on, high-ownership work: you'll build and operate real infrastructure, not manage tickets.

What You Will Do
  • Own reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
  • Perform hands-on infrastructure work - server provisioning, OS configuration, networking, storage, and hardware troubleshooting - from bare metal through auto-scaling Kubernetes environments.
  • Lead capacity planning and hardware lifecycle management for your domains, and track cloud spend to support FinOps and workload placement decisions.
  • Drive all provisioning, deployment, and operational changes through Terraform and/or Ansible rather than manual steps, and contribute to shared IaC modules used across the global SRE and data center services teams.
  • Build and document automation that eliminates toil - host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.
  • Design and maintain monitoring, alerting, and SLIs (Prometheus/Grafana, DataDog, Splunk, or equivalent), contributing to AIOps-driven detection workflows.
  • Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce high-quality RCAs for P0/P1 incidents.
  • Support platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.

What You Will Bring
  • Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.
  • Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure - networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.
  • Production IaC experience with Terraform and/or Ansible - writing and maintaining configurations, not just running existing playbooks.
  • Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.
  • Experience with observability tooling - Prometheus/Grafana, DataDog, Splunk, or equivalent - including building dashboards and writing alert rules.
  • Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.

Preferred Qualifications
  • Experience operating customer-facing infrastructure or platform services with external reliability expectations.
  • Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.
  • Experience deploying and operating AI-driven infrastructure tools - AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics - in production.
  • HPC job scheduler experience: Slurm, LSF, or equivalent.
  • Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.
  • Experience with large-scale infrastructure automation - host lifecycle management, fleet auto-healing, or AIOps-driven operations - building tooling that reduces manual intervention, not just running it.

Similar Jobs

More Jobs at d-Matrix

More Enterprise Technology Jobs

Find similar Sr Staff Site Reliability Engineer, AI Infrastructure jobs: