ServiceNow

Director, Data & Storage Reliability Engineering

ServiceNow$221K — $387K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 15+ years of experience in software or platform engineering, including SaaS environments
  • 8+ years of engineering leadership experience
  • Deep expertise in distributed systems, cloud infrastructure, and database technologies
  • Experience defining product strategies and managing engineering portfolios
  • Strong understanding of reliability engineering principles and operational excellence
  • Exceptional communication and stakeholder management skills
  • Bachelor's degree in Computer Science, Engineering, or a related field.

Responsibilities

  • Lead and scale high-performing engineering teams focused on reliability and performance
  • Drive the establishment of engineering practices and standards across the organization
  • Evaluate incidents to identify systemic risks and drive engineering improvements
  • Partner with cross-functional teams to embed reliability considerations in software development
  • Transform production insights into strategic engineering initiatives
  • Champion a predictive reliability engineering model that emphasizes prevention and continuous optimization
  • Develop meaningful KPIs to measure reliability and performance.

Benefits

  • Health plans including flexible spending accounts
  • 401(k) Plan with company match
  • Employee stock purchase plan
  • Matching donations program
  • Flexible time away and family leave programs
  • Work persona flexibility allowing remote or in-office work based on individual roles.
Full Job Description
Job Description

What you get to do in this role:

Team Management

The successful candidate will lead the Data & Storage Reliability Engineering organization responsible for improving reliability, resilience, performance, scalability, observability, and customer experience across ServiceNow's database, storage, and supporting platform infrastructure.

This leader will be responsible for building, developing, and scaling high-performing engineering teams focused on reliability engineering, observability, performance engineering, diagnostics, automation, production analytics, migration readiness, resilience engineering, and prevention engineering.

Responsibilities include talent acquisition, performance management, career development, succession planning, objective setting, coaching, and prioritization of strategic initiatives.

The role will establish a strong engineering-first culture centered on data-driven decision making, continuous improvement, operational excellence, customer experience, and systemic risk reduction.

This position is accountable for identifying recurring failure patterns, reliability risks, performance bottlenecks, scalability constraints, migration challenges, and operational inefficiencies across database services, storage platforms, cloud infrastructure, and distributed application environments, and driving engineering improvements that eliminate entire classes of issues before they impact customers.

The Director will partner closely with Product Engineering, Database Engineering, Cloud Infrastructure, Architecture, Storage Engineering, Support, and Operations teams to ensure reliability, observability, performance, and resilience considerations are incorporated throughout the software development lifecycle.

The successful candidate will also partner closely with SWAT and Customer & Production Engineering teams to establish a continuous feedback loop between production operations and platform improvement. SWAT remains responsible for customer escalations, production operations, incident response, and service restoration, while this organization is responsible for identifying systemic opportunities, defining engineering priorities, and driving platform improvements that reduce future customer impact.

The successful candidate will serve as the senior technical leader for complex reliability investigations, customer-critical escalation reviews, migration readiness assessments, and platform improvement initiatives, transforming production insights into long-term engineering outcomes.

They will influence architectural decisions and technology investments by providing reliability expertise, observability insights, performance guidance, and production-based evidence that improve platform resilience, scalability, efficiency, and customer outcomes.

This role requires a strong product mindset. The leader will treat reliability, observability, resilience, performance, and automation capabilities as products with roadmaps, priorities, adoption goals, and measurable outcomes. They will be responsible for identifying the highest-value engineering opportunities, prioritizing investments, and driving adoption across multiple product and infrastructure organizations.

Process and Procedures

The successful candidate will establish scalable reliability engineering practices, standards, governance processes, and operating models across the organization.

They will drive adoption of observability standards, reliability engineering frameworks, resiliency assessments, migration readiness practices, diagnostics capabilities, engineering guardrails, and automation strategies.

This leader will continuously evaluate incidents, customer escalations, migration outcomes, platform telemetry, performance trends, capacity signals, and operational data to identify systemic risks and drive long-term engineering improvements.

The role will establish a formal review process with SWAT and Customer & Production Engineering teams to evaluate major incidents, recurring operational challenges, migration learnings, customer-impacting events, and emerging platform risks. These insights will be used to prioritize engineering investments and platform improvements.

The successful candidate will establish meaningful KPIs and engineering metrics that provide visibility into platform reliability, resiliency, performance, operational efficiency, customer experience, engineering productivity, and risk reduction.

The successful candidate will leverage AI-powered tools, analytics, automation frameworks, and production intelligence to identify emerging risks, improve detection coverage, accelerate engineering insights, reduce operational toil, and improve engineering productivity.

They will use production telemetry, incident learnings, customer escalations, migration outcomes, observability data, and operational trends to drive architectural improvements, reliability investments, platform standards, and long-term engineering evolution.

The Director will maintain a portfolio of reliability investments spanning observability, performance, diagnostics, resilience, automation, and prevention, balancing immediate customer needs with long-term platform strategy.

The Director will champion a proactive reliability engineering model that shifts the organization from reactive issue response toward predictive analysis, prevention, resilience, and continuous optimization.

Qualifications

To be successful in this role you have:
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • Strong product mindset with demonstrated experience treating technical capabilities as products with roadmaps, priorities, customers, adoption goals, and measurable business outcomes.
  • Experience translating production insights, customer pain points, operational challenges, reliability risks, and platform telemetry into prioritized engineering investments and long-term roadmaps.
  • Experience partnering closely with production operations, customer escalation teams, reliability organizations, and software engineering teams to drive systemic improvements based on operational learnings.
  • Experience defining product strategies, developing roadmaps, prioritizing investments, and aligning stakeholders across multiple organizations without direct authority.
  • Experience operating a portfolio of engineering investments, balancing short-term customer needs with long-term reliability, performance, scalability, and resilience objectives.
  • 15+ years of experience in software engineering, platform engineering, reliability engineering, infrastructure engineering, database engineering, distributed systems, product management, or large-scale SaaS environments.
  • 8+ years of engineering leadership experience, including leading managers and globally distributed teams.
  • Extensive experience leading Reliability Engineering, Platform Engineering, Database Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical organizations.
  • Deep expertise in distributed systems, databases, storage technologies, cloud infrastructure, and large-scale SaaS architectures.
  • Strong understanding of reliability engineering principles, observability, scalability, resiliency, operational excellence, and performance engineering.
  • Experience building and operating observability, telemetry, diagnostics, reliability, or performance capabilities at scale.
  • Proven experience identifying systemic issues and converting operational insights into strategic engineering improvements.
  • Experience partnering closely with Product Management organizations to influence roadmaps and deliver customer-centric outcomes.
  • Experience driving engineering initiatives through data, metrics, customer impact analysis, and measurable business outcomes.
  • Experience leveraging AI technologies to improve decision-making, analytics, engineering workflows, operational efficiency, reliability insights, automation, or customer outcomes.
  • Exceptional communication, stakeholder management, and leadership skills.
  • Bachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.

Desired Skills
  • Previous Product Management experience in a platform, infrastructure, cloud, database, storage, or SaaS environment.
  • Experience applying product management disciplines such as roadmap planning, prioritization, customer-centric thinking, outcome measurement, and portfolio management to engineering organizations.
  • Experience operating large-scale enterprise database and storage platforms supporting mission-critical workloads.
  • Experience building and scaling Reliability Engineering, Performance Engineering, Platform Engineering, SRE, or Production Engineering organizations.
  • Experience with observability platforms, telemetry systems, diagnostics frameworks, and production analytics.
  • Experience with migration readiness, resiliency validation, reliability testing, operational risk reduction, and large-scale cloud transformations.
  • Experience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, and engineering productivity.
  • Strong understanding of distributed systems architecture, cloud platform operations, and hyperscale environments.
  • Experience developing executive-facing reliability scorecards, engineering metrics, and business impact reporting.
  • Experience influencing platform architecture, database strategy, storage strategy, and long-term engineering roadmaps.
  • Experience with Linux-based production environments and large-scale cloud infrastructure.
  • Experience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms.
  • Familiarity with ServiceNow platform architecture and large-scale SaaS operations.


JV20

For positions in this location, we offer a base pay of $221,200 - $387,100, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

About ServiceNow

ServiceNow provides cloud-based solutions that define, structure, manage, and automate services for enterprise operations in North America, Europe, the Middle East, Africa, the Asia Pacific, and other countries. The company offers service management solutions, including incident, problem, change, request, and cost management as well as service catalogs; and IT, HR, facilities, and field service management solutions. It also provides IT operations management solutions covering service mapping, delivery, and assurance solutions; business management solutions such as financial management, project portfolio suite, vendor performance management, and performance analytics as well as governance, risk, and compliance; and application development services.

ServiceNow Careers

Join the dynamic team at ServiceNow, a global leader in digital workflow solutions, where innovation and leadership converge to shape the future of work. At ServiceNow, we offer more than just job opportunities; we provide a platform for professional growth and a chance to be part of a culture that values diversity, creativity, and continuous learning.

Work You’ll Do

Embark on a career journey with ServiceNow and contribute to the world’s leading enterprises' digital transformation. Our team is at the forefront of developing cutting-edge technologies that improve how people work. With ServiceNow, you will use your skills to impact businesses and industries profoundly, driving efficiency and innovation.

Join Our Market-Leading Team

ServiceNow is not just another technology company. We are a team that thrives on diversity and leadership, fostering an inclusive environment that promotes growth and development. Our commitment to diversity training ensures that every team member can achieve their potential.

Innovative Work

ServiceNow is home to more than 10,000 dedicated professionals who lead the charge in digital workflows and enterprise solutions. As part of our team, you will engage in projects that merge technology with practical applications, creating revolutionary products that advance how services are delivered and managed.

Career Development

At ServiceNow, your career trajectory is filled with boundless opportunities. We support your growth with robust training programs, leadership development courses, and access to global challenges. Whether you are looking for an internship, full-time position, or leadership role, ServiceNow equips you with the tools to excel.

Be Part of a Great Team

Working at ServiceNow means being part of a community that values teamwork and innovation. Our collaborative environment encourages networking and sharing ideas, making our workplace vibrant and dynamic. The benefits of joining ServiceNow extend beyond comprehensive health and wellness; they include fostering professional connections and friendships that last a lifetime.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional or a recent graduate, ServiceNow offers a range of employment options to suit your career goals. From internships that provide real-world experience to full-time positions that challenge you to leverage your expertise, we are committed to hiring the best talent.

Stay Connected

Join Our Team Search open positions that match your skills and interests. At ServiceNow, we look for passionate, curious, and solution-driven team players. Explore the possibilities that await you at a company that is committed to your professional success.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at ServiceNow.

ServiceNow Careers

Empowering professionals to achieve more, ServiceNow is where careers are future-proofed, and ambitions are realized. Join us in our journey of growth and innovation.
Learn more about ServiceNow
Size
16,881 employees
Market Cap
$76.5 billion
Industry
Net Income
$118.5 million
Founded
2004
5 Year Trend
+33.5%
Revenue
$4.5 billion
NASDAQ

Similar Jobs

More Jobs at ServiceNow

More Information Technology Jobs

Find similar Director, Data & Storage Reliability Engineering jobs: