SAP

Expert Reliability Engineer

SAP$176K — $374K *
Technical Services
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years in DevOps/SRE engineering
  • 7+ years leading engineering projects
  • Deep expertise in SRE principles for mission-critical platforms
  • Advanced experience in monitoring tools like Grafana and Prometheus
  • Strong background in security roles and practices
  • Expert knowledge in Linux/Unix and large-scale distributed environments
  • Proficient in IaC and automation tools like Terraform, Ansible, and scripting languages.

Responsibilities

  • Secure and scale the foundational platform for SAP Sovereign Cloud
  • Lead a globally distributed team in setting best practices
  • Manage operational reliability and platform security standards
  • Conduct code reviews and enhance AI-assisted workflows
  • Implement backup and disaster recovery drills
  • Identify and address capability gaps in operational processes
  • Collaborate closely with regional teams and hyperscaler partners.

Benefits

  • Hybrid work model with flexibility
  • Opportunity to lead innovative SRE initiatives
  • Access to cutting-edge AI technologies
  • Mentorship and growth opportunities in a global team
  • Engagement with a mission-critical cloud platform.
Full Job Description
***** Due to the potentially classified nature of our work, your willingness is required to subject yourself to a governmental security clearance process ******

YOUR FUTURE ROLE

We are looking for an Expert Reliability Engineer (SRE) within the Shared Management Services (SMS) group in the Technology and Engineering unit of SAP Sovereign Cloud organization.

In this role, you will join the Technology and Engineering team as a Site Reliability Engineer focused on securing and scaling the foundational platform that underpins SAP Sovereign Cloud. You will work alongside a globally distributed team of highly motivated engineers responsible for the design, development, deployment, and lifecycle management of the Sovereign Cloud Shared Management Services (SMS) platform, with a mandate that spans both operational reliability and platform security.

As an expert Site Reliability Engineer, you will shoulder a shared ownership of the reliability, security, and operational excellence of production and non-production environments within Sovereign Cloud. It is expected that you will have extensive experience in all areas of a critical application administration stack spanning:

  • source control (git)
  • CI/CD platforms
  • identity and access management
  • secrets management
  • container orchestration
  • network security infrastructure
  • full-stack observability tooling across multi-cloud environments.
  • AI-assisted engineering workflows


You will lead an experienced team of globally distributed engineers in setting standards and best practices across responsibility areas including:
  • Code reviews
  • Agentic coding (AI) security best practices
  • Agentic coding workflows and skill development
  • AI-assisted incident response
  • Unit test coverage
  • Functional test coverage
  • etc


You will treat security as a first-class reliability concern: hardening identity and access management, secrets management, and supply chain integrity are as central to this role as uptime and incident response.

You will identify and close mission-critical capability gaps, define disciplined and standardized operational processes, and help the team navigate trade-offs across deployment plans, infrastructure investments, and day-to-day operational decisions. You will own and continuously improve backup and disaster recovery drills, ensuring the platform is failure-ready at global scale.

You will work closely with Sovereign Cloud operations and engineering teams, regional counterparts, and hyperscaler provider partners.

WHAT YOU BRING
  • Ability to manage ambiguities while being innovative and collaborative
  • Extensive technology skills and the willingness to learn new topics quickly
  • Problem-solving, presentation, communication, and interpersonal skills
  • Ability to think strategically, delivering projects and work cross-organizationally
  • Knowledge of SAP and the SAP solution portfolio
  • Cultural awareness, intercultural competencies, and the ability to influence without formal authority
  • Ability to build trusted relationships with key stakeholders
  • Persistence, self-motivation, and willingness to work under pressure
  • Proven ability to work in cross-functional teams
  • Ability to lead and mentor junior engineers in setting and maintaining DevOps and SRE best practices
  • English (fluent)


WORK EXPERIENCE
  • 10+ years of experience in DevOps and/or SRE engineering
  • 7+ years of experience with a successful track record of leading engineering projects and cross-functional program teams
  • Deep mastery of SRE principles as they apply to a globally distributed, mission-critical platforms
  • Advanced Experience in architecture, engineering, and deployment of modern monitoring tooling such as Grafana, Promethius
  • Background in security or security-adjacent roles with a proven record of strong security foundational skills
  • Proven track record in end-to-end implementation initiatives
  • Experience in people management or staff level technical leadership
  • Expert level knowledge of Linux/Unix administration, networking fundamentals, and operating large-scale distributed cloud environments
  • Advanced proficiency with IaC, scripting, and automation such as Ansible, Terraform, Bash, Python, Powershell, CI/CD
  • Deep expertise with modern access control, authentication standards and identity federation at an enterprise scale.
  • Proven experience architecting AI-driven incident response workflows that reduce mean time to detection and resolution through automated anomaly correlation, intelligent alert triage, and LLM-assisted runbook execution
  • Proven ability to own the design, implementation, and continuous refinement of SRE reliability frameworks, including SLAs, SLOs, and SLIs, ensuring alignment between platform performance, business commitments, and engineering priorities

Requisition ID: 458145 | Work Area: Software-Development Operations | Expected Travel: 0 - 10% | Career Status: Professional | Employment Type: Regular Full Time | Additional Locations: #LI-Hybrid

Requisition ID: 458145

Posted Date: Sep 1, 2026

Work Area: Software-Development Operations

Career Status: Professional

Employment Type: Regular Full Time

Expected Travel: 0 - 10%

Location:

Similar Jobs

More Jobs at SAP

More Technical Services Jobs

Find similar Expert Reliability Engineer jobs: