Principal, Staff Site Reliability

DigitalBridge Group, Inc.

• $150K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in building and operating production infrastructure at scale, including hybrid cloud and on-prem environments.
  • Deep expertise in AWS and Azure services, including compute, networking, IAM, and managed data services.
  • Expert-level skills in Linux, Terraform, Ansible, and modern CI/CD tools like GitHub Actions or GitLab CI.
  • Proven track record of improving platform reliability, MTTR, deployment velocity, and developer productivity.
  • Strong scripting skills in Python and/or Go, with the ability to debug across the stack.
  • Production experience with Kubernetes, service meshes, and container security posture.
  • Fluency with observability tools such as Datadog, Prometheus, and ELK/Splunk.

Responsibilities

  • Own end-to-end reliability for business-critical services, focusing on SLOs, error budgets, and incident command.
  • Design and evolve multi-cloud and on-prem infrastructure, driving workload placement and resilience trade-offs.
  • Build and maintain the CI/CD backbone using Terraform and Ansible to enable safe and quick deployments.
  • Advance observability to help on-call engineers localize failures rapidly and instrument reliability as a product surface.
  • Lead the development of AI-assisted SRE capabilities, including LLM-driven incident response and automated RCA.
  • Collaborate with security and application teams on hardening and change management in a regulated environment.
  • Participate in on-call rotation and conduct blameless postmortems to drive systemic fixes.

Benefits

  • Opportunity to influence architecture and tooling standards in a senior IC role.
  • Hands-on experience with cutting-edge AI technologies in incident response.
  • Mentorship opportunities for senior engineers, fostering a culture of learning.
  • Engagement in a regulated environment, enhancing skills in compliance and security practices.
Full Job Description
We are hiring a Staff SRE / Infrastructure Engineer to own the reliability, scalability, and developer experience of the platforms that power our investment, portfolio, and enterprise operations. You will set the technical direction for our hybrid cloud footprint (AWS + Azure +

on-prem), harden production for a regulated buy-side environment, and lead an emerging body of work applying AI to incident response and root-cause analysis. This is a hands-on senior IC role with material influence over architecture, tooling standards, and how engineers ship software.

What you'll do
  • Own end-to-end reliability for business-critical services - SLOs, error budgets, capacity planning, DR, and incident command - with a five-nines mindset appropriate to financial services workloads.
  • Design and evolve our multi-cloud and on-prem infrastructure across AWS, Azure, and colocated environments; drive workload placement, cost, and resilience trade-offs.
  • Build and maintain the Terraform, Ansible, and CI/CD backbone that lets product and data teams ship safely and quickly; codify golden paths and paved roads.
  • Advance observability (metrics, logs, traces, profiling) so on-call engineers can localize failures in minutes, not hours; instrument reliability as a first-class product surface.
  • Lead the buildout of AI-assisted SRE capabilities - LLM-driven triage, incident summarization, runbook synthesis, and automated RCA - with human-in-the-loop guardrails.
  • Partner with security, data platform, and application teams on hardening, patching, secrets, network segmentation, and change management appropriate to a regulated environment.
  • Participate in a leader-level on-call rotation; run blameless postmortems and drive systemic fixes to closure.
  • Mentor senior engineers; set the bar for infrastructure code review, production readiness reviews, and reliability practice across the org.


Required experience
  • 10+ years building and operating production infrastructure at scale, including hybrid cloud

+ on-prem.
  • Deep expertise in AWS and Azure compute, networking, IAM, and managed data services; comfortable with account/project topology, landing zones, and org-level guardrails.
  • Expert-level Linux, Terraform, Ansible, and modern CI/CD (GitHub Actions, GitLab CI, Argo, or equivalent).
  • Track record of measurable improvements in platform reliability, MTTR, deployment velocity, and developer productivity.
  • Strong scripting/software skills in Python and/or Go; can read and refactor application code well enough to debug across the stack.
  • Production experience with Kubernetes, service meshes, and container security posture.
  • Fluency with observability stacks (Datadog, Prometheus/Grafana, OpenTelemetry, ELK/Splunk).
  • Experience running incident command and driving durable postmortem outcomes.


Nice to have
  • Prior experience in financial services, buy-side, or another regulated environment (SOC 2, SOX, GLBA).
  • Hands-on work with AI/LLM tooling for SRE - agentic incident response, RAG over runbooks/telemetry, or automated RCA.
  • Experience with data-center automation, colocation, or modular infrastructure.
  • FinOps depth: unit economics, cost attribution, and budget governance across AWS/Azure.


Similar Jobs

More Jobs at DigitalBridge Group, Inc.

More Information Technology Jobs

Find similar Principal, Staff Site Reliability jobs: