To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.
Job Category
Software Engineering
Job Details
Job Title: Director, Site Reliability Engineering
Location: New York, NY; San Francisco, CA; Dallas, TXAbout the RoleWe are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.
In this role, you will transform our SRE function-moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.
As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.
Key ResponsibilitiesSRE Strategy and Leadership- Define and execute the long-term strategy and roadmap for Site Reliability Engineering.
- Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.
- Build and develop a high-performing team of site reliability and operations engineers.
- Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.
- Translate business priorities and customer impact into clear reliability investments and engineering outcomes.
- Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.
Reliability Engineering- Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.
- Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.
- Define what it means for a service to be operationally and observably ready for production.
- Develop readiness reviews and certification practices for high-impact services and launches.
- Drive improvements in system availability, performance, resiliency, and recovery.
- Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.
Observability- Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.
- Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.
- Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.
- Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.
- Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.
- Establish governance and measurement to assess adoption and effectiveness of observability standards.
Incident Management and Operational Excellence- Improve incident detection, response, mitigation, communication, and learning.
- Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.
- Reduce mean time to detect, acknowledge, mitigate, and recover.
- Improve on-call practices, escalation paths, runbooks, and operational ownership.
- Establish blameless post-incident review practices that produce measurable engineering improvements.
- Identify recurring sources of operational toil and create plans to eliminate or automate them.
- Partner with engineering leaders to ensure actions from incidents are prioritized and completed.
Automation and AI-Enabled Operations- Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.
- Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.
- Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.
- Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.
- Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.
Cross-Functional Partnership- Partner with engineering, DevOps, and business stakeholders.
- Influence teams that do not directly report into SRE and build shared accountability for production outcomes.
- Create clear service ownership models and operational expectations across teams.
- Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.
- Communicate reliability posture, risks, trends, and investments to executive and technical audiences.
Measures of Success- Improved availability and reliability of critical services.
- Reduced time to detect, diagnose, mitigate, and recover from incidents.
- Increased percentage of services meeting observability and production-readiness standards.
- Reduced alert noise, operational toil, and manual incident-management activity.
- Increased adoption of service-level objectives and measurable reliability practices.
- Improved quality and completion rate of post-incident corrective actions.
- Increased automation across detection, triage, remediation, and operational workflows.
- Stronger ownership of production reliability across engineering teams.
Leadership Attributes- Engineering-first and automation-oriented.
- Comfortable challenging legacy operating models and assumptions.
- Able to move between technical detail and executive-level strategy.
- Outcome-focused, pragmatic, and data-driven.
- Builds trust through clarity, accountability, and strong partnership.
- Develops leaders and creates an inclusive, high-performance engineering culture.
- Treats incidents as opportunities to improve systems rather than assign blame.
- Brings urgency to operational risks while maintaining focus on sustainable solutions.
Minimum Qualifications- Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master's degree or MBA preferred.
- 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
- Proven experience building or transforming a reliability or operational engineering organization.
- Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.
- Experience establishing observability, incident-management, service-level objective, and production-readiness practices.
- Demonstrated ability to improve reliability through engineering and automation rather than process alone.
- Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.
- Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.
- Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders
- Ability to balance immediate operational needs with long-term engineering transformation.
- Strong written, verbal, and executive communication skills.
Preferred Qualifications- Strong experience operating large-scale systems in AWS or another major cloud environment.
- Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.
- Demonstrated experience implementing OpenTelemetry or common instrumentation standards.
- Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.
- Experience applying AI, machine learning, or agent-based automation to operational workflows.
- Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.
- Software engineering experience and the ability to engage deeply in architecture and design discussions.
- Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.
*LI-Y