Qualifications
Responsibilities
Benefits
Job Title CaaS Private Site Reliability Engineer
Corporate Title Assistant Vice President
Location Cary, NC
Overview
As a Site Reliability Engineer on the CaaS Private platform team, you will help operate and improve an on-prem, multi-tenant Kubernetes platform running on bare metal. You will strengthen the reliability, observability, scalability, and operational excellence of a platform that supports critical, low-latency, and regulated workloads. You will partner closely with platform, network, security, and application teams to define service level objectives, improve resilience, reduce operational toil, and build automation that allows the platform to run safely at scale. Join us here, and you will turn operational challenges into measurable engineering improvements that application teams can rely on every day.
What We Offer You
A diverse and inclusive environment that embraces change, innovation, and collaboration
A hybrid working model, allowing for in-office / work from home flexibility, generous vacation, personal and volunteer days
Employee Resource Groups support an inclusive workplace for everyone and promote community engagement
Competitive compensation packages including health and wellbeing benefits, retirement savings plans, parental leave, and family building benefits
Educational resources, matching gift and volunteer programs
What You’ll Do
Define, implement, and continuously improve SLI/SLOs, alerting standards, and error budgets for the CaaS Private platform and critical services
Build and maintain observability across metrics, logs, alerts, and dashboards to provide clear insight into platform health, saturation, latency, and failure modes
Lead or coordinate incident response for platform-impacting events, ensuring timely mitigation, clear communication, blameless postmortems, and durable follow-up actions
Automate repetitive operational tasks and remediation workflows to reduce toil, improve platform consistency, and accelerate recovery time
Improve reliability, upgrade safety, and operational readiness for Kubernetes clusters, ingress paths, service mesh components, node services, and critical platform dependencies
Partner with platform, network, security, and application teams on capacity planning, release readiness, troubleshooting, operational documentation, and adoption of best practices
Skills You’ll Need
Proven experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure role
Hands-on Kubernetes expertise, including operating clusters on bare metal or private cloud environments and supporting platform services at scale
Strong Linux system administration capability and infrastructure-level scripting experience using Python, Ansible, and Bash
Practical knowledge of observability stacks and telemetry pipelines, including Prometheus, Grafana, Splunk, metrics, logging, alerting, dashboards, and OpenTelemetry-style concepts
Strong understanding of incident management, root cause analysis, operational readiness, computer networking, virtualization, containerization, and distributed systems behavior under failure
Skills That Will Help You Excel
Experience with Istio / Envoy, service mesh observability, traffic management, OPA Gatekeeper, admission controls, or policy-driven operational guardrails
Familiarity supporting stateful services such as PostgreSQL, Kafka, MongoDB, or comparable platform dependencies
Practical knowledge of capacity planning, load testing, chaos testing, failure-injection techniques, alert tuning, and self-healing automation
Exposure to low-latency or regulated environments with strict uptime, change control, compliance, time synchronization, deterministic performance, or SR-IOV workload constraints
Ability to read and understand Golang code when troubleshooting platform components, with strong written and verbal communication skills and a continuous learning mindset
Expectations
It is the Bank’s expectation that employees hired into this role will work in the Cary, NC office in accordance with the Bank’s hybrid working model.
The salary range for this position in Cary is $100,000 to $153,000.Actual salaries may be based on a number of factors including, but not limited to, a candidate’s skill set, experience, education, work location and other qualifications. Posted salary ranges do not include incentive compensation or any other type of remuneration.
Deutsche Bank Benefits
At Deutsche Bank, we recognize that our benefit programs have a profound impact on our colleagues. That’s why we are focused on providing benefits and perks that enable our colleagues to live authentically and be their whole selves, at every stage of life. We provide access to physical, emotional, and financial wellness benefits that allow our colleagues to stay financially secure and strike balance between work and home. Click to learn more!
Learn more about your life at Deutsche Bank through the eyes of our current employees:
#LI-HYBRID
About Deutsche Bank
Similar Jobs



More Jobs at Deutsche Bank




More Information Technology Jobs