Deutsche Bank

CaaS Private Site Reliability Engineer - Assistant Vice President

Deutsche Bank$100K — $153K *
Cary, NC 27513In-Person
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in Site Reliability Engineering, Production Engineering, or DevOps.
  • Proven hands-on expertise with Kubernetes in bare metal environments.
  • Strong Linux system administration and scripting skills in Python, Ansible, and Bash.
  • Knowledge of observability tools like Prometheus, Grafana, and OpenTelemetry.
  • Understanding of incident management, root cause analysis, and distributed systems.

Responsibilities

  • Define and improve SLI/SLOs and alerting standards for the platform.
  • Build observability through metrics, logs, and dashboards to monitor platform health.
  • Lead incident response for platform issues, ensuring timely communication and follow-ups.
  • Automate operational tasks to enhance platform consistency and reduce manual effort.
  • Improve reliability and operational readiness for critical dependencies and services.

Benefits

  • Diverse and inclusive work environment fostering collaboration and innovation.
  • Hybrid working model with flexibility for in-office and remote work.
  • Generous vacation and personal leave, including volunteer days.
  • Health, wellness, and family building benefits included in competitive packages.
  • Access to educational resources and employee resource groups for community engagement.
Full Job Description
Job Description:

Job Title CaaS Private Site Reliability Engineer

Corporate Title Assistant Vice President

Location Cary, NC

Overview

As a Site Reliability Engineer on the CaaS Private platform team, you will help operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal. You will strengthen the reliability, observability, scalability, and operational excellence of a platform that supports critical, low-latency, and regulated workloads. You will partner closely with platform, network, security, and application teams to define service level objectives, improve resilience, reduce operational toil, and build automation that allows the platform to run safely at scale. Join us here, and you will turn operational challenges into measurable engineering improvements that application teams can rely on every day.

What We Offer You

  • A diverse and inclusive environment that embraces change, innovation, and collaboration

  • A hybrid working model, allowing for in-office / work from home flexibility, generous vacation, personal and volunteer days

  • Employee Resource Groups support an inclusive workplace for everyone and promote community engagement

  • Competitive compensation packages including health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits

  • Educational resources, matching gift and volunteer programs

What You’ll Do

  • Define, implement, and continuously improve SLI/SLOs, alerting standards, and error budgets for the CaaS Private platform and critical services

  • Build and maintain observability across metrics, logs, alerts, and dashboards to provide clear insight into platform health, saturation, latency, and failure modes

  • Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear communication, blameless postmortems, and durable follow-up actions

  • Automate repetitive operational tasks and remediation workflows to reduce toil, improve platform consistency, and accelerate recovery time

  • Improve reliability, upgrade safety, and operational readiness for Kubernetes clusters, ingress paths, service mesh components, node services, and critical platform dependencies

  • Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, operational documentation, and adoption of best practices

Skills You’ll Need

  • Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure role

  • Hands-on Kubernetes expertise, including operating clusters on bare metal or private cloud environments and supporting platform services at scale

  • Strong Linux system administration capability and infrastructure-level scripting experience using Python, Ansible, and Bash

  • Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, dashboards, and OpenTelemetry-style concepts

  • Strong understanding of incident management, root cause analysis, operational readiness, computer networking, virtualization, containerization, and distributed systems behavior under failure

Skills That Will Help You Excel

  • Experience with Istio / Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails

  • Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies

  • Practical knowledge of capacity planning, load testing, chaos testing, failure-injection techniques, alert tuning, and self-healing automation

  • Exposure to low-latency or regulated environments with strict uptime, change control, compliance, time synchronization, deterministic performance, or SR-IOV workload constraints

  • Ability to read and understand Golang code when troubleshooting platform components, with strong written and verbal communication skills and a continuous learning mindset

Expectations

It is the Bank’s expectation that employees hired into this role will work in the Cary, NC office in accordance with the Bank’s hybrid working model.

The salary range for this position in Cary is $100,000 to $153,000.Actual salaries may be based on a number of factors including, but not limited to, a candidate’s skill set, experience, education, work location and other qualifications. Posted salary ranges do not include incentive compensation or any other type of remuneration.

Deutsche Bank Benefits

At Deutsche Bank, we recognize that our benefit programs have a profound impact on our colleagues. That’s why we are focused on providing benefits and perks that enable our colleagues to live authentically and be their whole selves, at every stage of life. We provide access to physical, emotional, and financial wellness benefits that allow our colleagues to stay financially secure and strike balance between work and home. Click to learn more!

Learn more about your life at Deutsche Bank through the eyes of our current employees:

#LI-HYBRID

About Deutsche Bank

Deutsche Bank AG is a German multinational investment bank and financial services company headquartered in Frankfurt, Germany. The bank is operational in 58 countries with a large presence in Europe, the Americas, and Asia. As of 2021, Deutsche Bank is the 21st largest bank in the world by total assets. The bank offers financial products and services for corporate and institutional clients along with private and business clients. Its services include sales, trading, research, and origination of debt and equity; mergers and acquisitions (M&A); risk management products, such as derivatives, corporate finance, wealth management, retail banking, fund management, and transaction banking.
Learn more about Deutsche Bank
Size
83,000 employees
Market Cap
$23.3 billion
Industry
Founded
1870
5 Year Trend
-8%
NASDAQ

Similar Jobs

More Jobs at Deutsche Bank

More Information Technology Jobs

Find similar CaaS Private Site Reliability Engineer - Assistant Vice President jobs: