Delinea

Manager, Site Reliability Engineering

Delinea$125K — $150K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations with ownership of production SaaS systems.
  • 2+ years of people leadership experience, including managing contractors.
  • Hands-on experience with Azure Kubernetes Service and AWS services including SES, EC2, RDS.
  • Deep observability expertise including metrics, logs, traces, and alerting strategies using Datadog or similar platforms.
  • Strong cloud networking fundamentals such as load balancing and identity management.
  • Automation skills in PowerShell, Python, or Bash, with infrastructure-as-code experience (Terraform, ARM, or Bicep).
  • Willingness to work across time zones and participate in an on-call rotation.

Responsibilities

  • Lead hands-on by actively engaging in code reviews, pipeline changes, and troubleshooting production issues.
  • Own availability and performance of Delinea Platform production environments across Azure and AWS.
  • Manage and develop a blended team of full-time SRE engineers and contractors.
  • Conduct meetings and planning sessions accommodating team members in different time zones.
  • Participate in on-call duties, acting as incident commander for critical events and ensuring effective communication during incidents.
  • Oversee end-to-end incident response, including post-incident review and root cause analysis.
  • Enhance observability by improving detection coverage and defining SLIs and SLOs.

Benefits

  • Opportunities for career growth and development.
  • Support for work-life balance through remote work and flexible hours.
  • Participation in a dynamic, distributed team environment.
  • Hands-on engagement with cutting-edge technology and tools.
  • Exposure to regulated industries with FedRAMP operations.
Full Job Description
Summary:

Delinea is looking for a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting the Delinea products. This is a working manager role. You will be expected to lead people and lead work: writing and reviewing automation, digging into AKS and Azure telemetry, commanding Sev1 and Sev2 incidents, and improving the observability and deployment practices your team depends on.

The initial scope is the Platform SRE and DevOps team. Over time, the role is expected to expand to cover our FedRAMP High environment and the broader Commercial Platform footprint, so comfort operating in a regulated environment and a willingness to grow scope are essential.

You will lead a blended team of full-time engineers and contractors distributed across multiple time zones. Meeting your team inside their working hours is an expectation of the role. On-call participation is required.

What You Will Do:
  • Lead hands-on. Spend a meaningful portion of your week in the environment: reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in Datadog, and troubleshooting production issues alongside your engineers. This role does not sit above the work.
  • Own availability and performance of the Delinea Platform production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF layers.
  • Manage a blended team. Hire, onboard, coach, and develop full-time SRE engineers. Direct and manage contractor resources, including scoping work, setting quality expectations, and reviewing deliverables.
  • Lead a distributed team. Run one-on-ones, standups, and planning sessions at times that work for engineers in other geographies.
  • Participate in on-call. Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, drive engagement of the right responders, and own communication cadence with support, engineering, and leadership until resolution.
  • Command incident response end to end. Own detection, triage, mitigation, customer-facing status communication, and post-incident review. Ensure RCAs are written to a customer-ready standard, preventative actions have owners and target dates, and those actions are driven to closure.
  • Raise the observability bar. Improve detection coverage so that issues are found by our monitoring rather than by a customer ticket. Own SLI and SLO definition, alert quality and noise reduction, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards.
  • Support FedRAMP and regulated operations. Grow into supporting our FedRAMP High environment, including change control discipline, evidence collection, boundary awareness, and the operational differences between government and commercial environments.
  • Reduce toil through automation. Set the expectation that repeat manual work becomes code. Prioritize automation backlog alongside project and reliability work.
  • Report on operational health. Produce and present incident metrics, trends, and reliability commitments to leadership, and translate them into a concrete improvement plan.

What You Will Need:
  • 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated ownership of production SaaS systems.
  • 2+ years of direct people leadership, including performance management, hiring, and coaching. Experience managing contractors or an outsourced delivery team is desired.
  • Current, hands-on production experience with the Delinea technology stack, including Azure Kubernetes Service, core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.
  • Hands-on experience across both Azure and AWS is required. You should be able to administer, troubleshoot, and reason about cost and security posture in each.
  • Deep observability expertise. Demonstrated ownership of an observability framework at scale: metrics, logs, traces, synthetics, SLOs, and alerting strategy. Hands-on proficiency with Datadog or an equivalent platform, including APM trace analysis and log-based troubleshooting.
  • Proven incident command. You have run major incidents as the incident commander, coordinated multiple responders under pressure, communicated to customers and executives during active impact, and authored the RCA afterward.
  • Strong cloud networking and security fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management.
  • Automation and scripting ability in PowerShell, Python, Bash, or similar, plus practical infrastructure-as-code experience (Terraform, ARM, or Bicep).
  • Practical experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery approaches.
  • Excellent written communication. You will write and approve customer-facing status updates and incident summaries under time pressure.
  • Willingness and availability to work across time zones and to participate in an on-call rotation.

We Would Love to See
  • Direct experience operating in a FedRAMP or other regulated environment (Azure Government, IL4/IL5, SOC 2, ISO 27001).
  • Experience standing up or maturing an incident management program, including sev definitions, escalation paths, on-call structure, and post-incident review process.
  • Experience with public status page operations and customer notification practices.
  • Experience with Atlassian Jira Service Management, Confluence, and Azure DevOps as the operational toolchain.
  • Track record of reducing customer-detected incidents through improved monitoring coverage.
  • Cost optimization experience across Azure and AWS on a meaningful scale.

For this Job, Delinea is not considering candidates that need any type of US work authorization now or in the future. This includes, but is not limited to: F1-OPT, F1-CPT, H-1B, TN, L-1, J1, etc.

About Delinea

Delinea is a pioneer in securing identities through centralized authorization, making organizations more secure by seamlessly governing their interactions across the modern enterprise. Delinea allows organizations to apply context and intelligence throughout the identity lifecycle across cloud and traditional infrastructure, data, and SaaS applications to eliminate identity-related threats. With intelligent authorization for all identities, Delinea is the only platform that enables you to identify each user, assign appropriate access levels, monitor interaction across the modern enterprise, and immediately respond upon detecting any irregularities. The Delinea Platform enables your teams to accelerate adoption and be more productive by deploying in weeks, not months, and requiring 10% of the resources to manage compared to the nearest competitor.
Learn more about Delinea
Size
500 employees
Industry
Founded
2004

Similar Jobs

More Jobs at Delinea

More Information Technology Jobs

Find similar Manager, Site Reliability Engineering jobs: