Salesforce

Director, Site Reliability Engineering

Salesforce$237K — $344K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or related field; Master's or MBA preferred
  • 10+ years of engineering experience, with 5+ years in leadership roles within SRE or Systems Engineering
  • Expertise in building and transforming reliability or operational engineering teams
  • Strong understanding of distributed systems and cloud architecture
  • Experience with establishing observability and incident management practices
  • Ability to improve reliability through engineering and automation
  • Proficient in modern telemetry tools and practices

Responsibilities

  • Define and execute a long-term strategy for Site Reliability Engineering
  • Establish operating models with clear scope, engagement, and success metrics
  • Build and develop a high-performing SRE team
  • Modernize SRE functions using automation and AI
  • Translate business priorities into engineering outcomes for reliability
  • Establish service-level standards and ensure production readiness

Benefits

  • Collaborative and inclusive work environment
  • Opportunity to drive engineering and reliability culture
  • Access to advanced technology tools for automation
  • Participation in high-impact projects with visibility across the organization
  • Commitment to professional development and leadership growth
Full Job Description
To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.

Job Category
Software Engineering

Job Details

Job Title: Director, Site Reliability Engineering
Location: New York, NY; San Francisco, CA; Dallas, TX


About the Role

We are looking for a Director of Site Reliability Engineering to spearhead the evolution of our reliability, observability, and operational engineering capabilities.

In this role, you will transform our SRE function-moving our engineering organization from reactive incident response to a proactive, automated, and data-driven reliability culture. Partnering closely across Application Engineering, Platform, Architecture, Security, Infrastructure, and Product, you will ensure our services are resilient, observable, scalable, and production-ready long before they launch.

As an impactful people leader with sharp technical judgment, you will directly manage and empower a core team of ~6 engineers while driving cross-functional alignment across a complex organization. You won't just run existing playbooks; you will define the strategy, tooling, automation, and culture needed to mentor your team and scale system reliability enterprise-wide.

Key Responsibilities

SRE Strategy and Leadership
  • Define and execute the long-term strategy and roadmap for Site Reliability Engineering.
  • Establish a clear operating model for SRE, including team scope, engagement models, ownership boundaries, and success measures.
  • Build and develop a high-performing team of site reliability and operations engineers.
  • Modernize the SRE function through automation, AI-assisted operations, self-service capabilities, and engineering-first practices.
  • Translate business priorities and customer impact into clear reliability investments and engineering outcomes.
  • Advise senior technology leaders on operational risk, resilience, capacity, and reliability tradeoffs.


Reliability Engineering
  • Establish service-level indicators, service-level objectives, error budgets, and reliability standards for critical services.
  • Partner with engineering teams to design reliability, scalability, recoverability, and graceful degradation into systems.
  • Define what it means for a service to be operationally and observably ready for production.
  • Develop readiness reviews and certification practices for high-impact services and launches.
  • Drive improvements in system availability, performance, resiliency, and recovery.
  • Ensure reliability requirements are incorporated throughout the software development lifecycle rather than addressed only after deployment.


Observability
  • Define an enterprise observability strategy spanning metrics, logs, traces, events, synthetics, real-user monitoring, and business telemetry.
  • Establish common instrumentation, telemetry, dashboards, alerting, and service-health standards.
  • Reduce fragmented or duplicative observability implementations by promoting shared patterns and reusable capabilities.
  • Improve end-to-end visibility across distributed systems, customer journeys, services, and infrastructure.
  • Partner with engineering teams to ensure telemetry is actionable, contextual, and tied to customer and business outcomes.
  • Establish governance and measurement to assess adoption and effectiveness of observability standards.


Incident Management and Operational Excellence
  • Improve incident detection, response, mitigation, communication, and learning.
  • Lead the transition from manual and reactive operations toward automated detection, diagnosis, remediation, and incident creation.
  • Reduce mean time to detect, acknowledge, mitigate, and recover.
  • Improve on-call practices, escalation paths, runbooks, and operational ownership.
  • Establish blameless post-incident review practices that produce measurable engineering improvements.
  • Identify recurring sources of operational toil and create plans to eliminate or automate them.
  • Partner with engineering leaders to ensure actions from incidents are prioritized and completed.


Automation and AI-Enabled Operations
  • Develop a roadmap for intelligent operations, including anomaly detection, event correlation, automated triage, assisted root-cause analysis, and remediation.
  • Evaluate opportunities to use agents and AI-assisted workflows across observability, incident response, capacity planning, and operational support.
  • Build automation that reduces cognitive load and improves the speed and consistency of operational decisions.
  • Ensure automation is safe, measurable, auditable, and designed with appropriate human oversight.
  • Promote platform and self-service approaches that allow product teams to adopt reliability practices with minimal friction.


Cross-Functional Partnership
  • Partner with engineering, DevOps, and business stakeholders.
  • Influence teams that do not directly report into SRE and build shared accountability for production outcomes.
  • Create clear service ownership models and operational expectations across teams.
  • Support major launches and critical business events through readiness planning, risk assessment, testing, and operational coordination.
  • Communicate reliability posture, risks, trends, and investments to executive and technical audiences.


Measures of Success
  • Improved availability and reliability of critical services.
  • Reduced time to detect, diagnose, mitigate, and recover from incidents.
  • Increased percentage of services meeting observability and production-readiness standards.
  • Reduced alert noise, operational toil, and manual incident-management activity.
  • Increased adoption of service-level objectives and measurable reliability practices.
  • Improved quality and completion rate of post-incident corrective actions.
  • Increased automation across detection, triage, remediation, and operational workflows.
  • Stronger ownership of production reliability across engineering teams.


Leadership Attributes
  • Engineering-first and automation-oriented.
  • Comfortable challenging legacy operating models and assumptions.
  • Able to move between technical detail and executive-level strategy.
  • Outcome-focused, pragmatic, and data-driven.
  • Builds trust through clarity, accountability, and strong partnership.
  • Develops leaders and creates an inclusive, high-performance engineering culture.
  • Treats incidents as opportunities to improve systems rather than assign blame.
  • Brings urgency to operational risks while maintaining focus on sustainable solutions.


Minimum Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Software Engineering, or a related technical field; Master's degree or MBA preferred.
  • 10+ years of progressive engineering experience, including 5+ years in engineering leadership managing SRE, Platform, or Systems Engineering teams
  • Proven experience building or transforming a reliability or operational engineering organization.
  • Strong understanding of distributed systems, cloud architecture, application architecture, networking, infrastructure, and software delivery.
  • Experience establishing observability, incident-management, service-level objective, and production-readiness practices.
  • Demonstrated ability to improve reliability through engineering and automation rather than process alone.
  • Proven experience leading teams responsible for highly available, customer-facing, or business-critical systems.
  • Strong understanding of modern telemetry, including metrics, logs, distributed tracing, synthetic monitoring, and real-user monitoring.
  • Demonstrated success driving alignment and building consensus across cross-functional engineering teams and executive stakeholders
  • Ability to balance immediate operational needs with long-term engineering transformation.
  • Strong written, verbal, and executive communication skills.

Preferred Qualifications
  • Strong experience operating large-scale systems in AWS or another major cloud environment.
  • Proven track record with observability platforms such as New Relic, Splunk, Datadog, Sentry, Honeycomb, Grafana, Prometheus, or OpenTelemetry.
  • Demonstrated experience implementing OpenTelemetry or common instrumentation standards.
  • Verified proficiency building internal developer platforms, paved roads, or self-service reliability capabilities.
  • Experience applying AI, machine learning, or agent-based automation to operational workflows.
  • Seasoned capability with chaos engineering, resilience testing, disaster recovery, capacity planning, and performance engineering.
  • Software engineering experience and the ability to engage deeply in architecture and design discussions.
  • Solid background in supporting high-profile launches, events, or systems with significant customer and business impact.


*LI-Y

About Salesforce

ExactTarget is a provider of on-demand email marketing software solutions. Their suite of on-demand one-to-one marketing applications enables clients to send business-critical and event-triggered communications to increase sales, optimize marketing investments, and strengthen customer relationships. They offer four editions of their on-demand software application along with integrated solutions such as ExactTarget for AppExchange and ExactTarget for [Microsoft](/organization/Microsoft) Dynamics CRM.

Salesforce Careers

Joining Salesforce means becoming part of a dynamic, global team of professionals who are deeply committed to driving customer success and innovation. As the world's leading Customer Relationship Management (CRM) platform, Salesforce offers unparalleled job opportunities in technology and consulting, making it an ideal place for ambitious individuals looking to make a significant impact.

Work You'll Do

At Salesforce, every position is a chance to leverage your skills and creativity to transform businesses and industries. Our diverse team of experts collaborates to deliver cutting-edge solutions that foster growth and enhance leadership capabilities. By joining our team, you'll be at the forefront of digital innovation, using Salesforce's powerful platform to help clients navigate their transformation journeys.

Innovate and Lead

Salesforce is not just a company; it's a community where you can lead with your ideas and see them come to life. Our culture of innovation encourages you to challenge the status quo and push the boundaries of what's possible. With Salesforce, you'll work alongside leaders in technology and business who are committed to your growth and professional development.

Career Growth and Opportunities

Whether you're looking for an internship, a full-time position, or leadership roles, Salesforce provides a wealth of opportunities to advance your career. Our commitment to professional growth is reflected in our robust training programs, including leadership development and diversity training, designed to help you excel at every stage of your career.

Be Part of a Great Team

Salesforce prides itself on a culture that values diversity, teamwork, and open communication. We believe that our strength lies in our people, and we're committed to creating an environment where everyone can thrive. Joining our team means being part of a supportive community that encourages networking and collaboration.

Benefits and Culture

At Salesforce, we understand that job satisfaction extends beyond the office. That's why we offer competitive benefits to support the health, well-being, and financial security of our employees and their families. From health insurance and retirement plans to wellness programs and flexible working arrangements, we provide the benefits that contribute to a better work-life balance.

Explore Job Opportunities

Ready to take the next step in your career? Explore the wide range of employment opportunities at Salesforce. From technical roles to customer engagement positions, we are continuously hiring talented individuals who are passionate about making a difference.

Stay Connected

Keep up to date with the latest at Salesforce by following our careers blog. Gain insights from the people who work here and learn how you can bring your career to the next level with Salesforce.

Apply Now

Are you ready to join a company that's leading the way in CRM technology? Search open positions that match your skills and interests on our careers page. Tailor your resume, prepare for your interview, and take the first step towards a rewarding career at Salesforce.

SEARCH SALESFORCE JOBS

Join Salesforce today and be part of a company that's shaping the future of technology, fostering a culture of innovation, and building a more equitable world.
Learn more about Salesforce
Size
73,541 employees
Market Cap
$130.4 billion
Industry
Net Income
$4 billion
Founded
2000
5 Year Trend
+25.7%
Revenue
$21.2 billion
NASDAQ

Similar Jobs

More Jobs at Salesforce

More Information Technology Jobs

Find similar Director, Site Reliability Engineering jobs: