Role Overviewd-Matrix's SRE team owns the infrastructure layer that every engineering team and customer depends on - colocation facilities, on-premises GPU clusters, cloud environments, and the platform services used to deploy and validate d-Matrix hardware and software. This role is a core member of that team, responsible for reliability, automation, and observability across colo, on-premises lab, and cloud environments. You will own systems end-to-end, from provisioning through live incident response, partnering with hardware and software teams on CI/CD, QA, and HPC workloads for silicon development, as well as supporting customer-facing environments where d-Matrix partners collaborate on deployments. This is hands-on, high-ownership work: you'll build and operate real infrastructure, not manage tickets.
What You Will Do- Own reliability and availability across colo server fleets, on-premises lab clusters, cloud environments (AWS, Azure, GCP), and customer-facing platform services.
- Perform hands-on infrastructure work - server provisioning, OS configuration, networking, storage, and hardware troubleshooting - from bare metal through auto-scaling Kubernetes environments.
- Lead capacity planning and hardware lifecycle management for your domains, and track cloud spend to support FinOps and workload placement decisions.
- Drive all provisioning, deployment, and operational changes through Terraform and/or Ansible rather than manual steps, and contribute to shared IaC modules used across the global SRE and data center services teams.
- Build and document automation that eliminates toil - host lifecycle management, fleet health checks, auto-remediation, self-service tooling, and networking automation for cluster interconnects and lab configurations.
- Design and maintain monitoring, alerting, and SLIs (Prometheus/Grafana, DataDog, Splunk, or equivalent), contributing to AIOps-driven detection workflows.
- Participate in on-call rotation, triaging and resolving incidents from bare metal to application layer, and produce high-quality RCAs for P0/P1 incidents.
- Support platform services used by internal teams and external customers, ensuring QoS and uptime commitments and documenting operational runbooks.
What You Will Bring- Bachelor's or Master's in Computer Science, Electrical Engineering, or a related field (or equivalent experience); 7+ years in SRE, infrastructure engineering, or systems administration.
- Strong Linux systems knowledge with hands-on colocation or on-premises server infrastructure - networking, storage, systemd, kernel parameters, performance diagnostics, physical hardware, rack networking, and bare-metal provisioning.
- Production IaC experience with Terraform and/or Ansible - writing and maintaining configurations, not just running existing playbooks.
- Kubernetes operational experience: cluster troubleshooting, workload management, storage, and networking.
- Experience with observability tooling - Prometheus/Grafana, DataDog, Splunk, or equivalent - including building dashboards and writing alert rules.
- Production-quality Python and/or Bash scripting, paired with incident response experience: structured triage, RCA production, and follow-through on action items.
Preferred Qualifications- Experience operating customer-facing infrastructure or platform services with external reliability expectations.
- Cloud infrastructure operations across AWS, Azure, or GCP, including hybrid environments spanning cloud and on-prem.
- Experience deploying and operating AI-driven infrastructure tools - AIOps platforms, intelligent alerting, anomaly detection, or LLM-assisted diagnostics - in production.
- HPC job scheduler experience: Slurm, LSF, or equivalent.
- Knowledge of high-speed interconnect fabrics: InfiniBand, RoCE, or NVLink.
- Experience with large-scale infrastructure automation - host lifecycle management, fleet auto-healing, or AIOps-driven operations - building tooling that reduces manual intervention, not just running it.