Delinea

Senior Site Reliability Engineer - FedRAMP

Delinea$130K — $155K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in Site Reliability Engineering, DevOps, cloud operations, or SaaS production engineering.
  • Hands-on experience with Azure services including AKS, App Service, Azure SQL.
  • Proficient in using Datadog for observability: metrics, logs, APM.
  • Experience with SLIs, SLOs, and error budgets in a production environment.
  • Incident response experience, including leading postmortems.
  • Proficient in Kubernetes and Terraform for infrastructure as code.
  • Scripting proficiency in PowerShell and Python.

Responsibilities

  • Own the reliability of production SaaS services, including availability and performance.
  • Define service-level indicators and set objectives to prioritize reliability work.
  • Build and optimize monitoring systems in Datadog and Azure Monitor.
  • Automate incident response and integrate monitors with remediation workflows.
  • Lead high-severity incident responses and write post-incident reviews.
  • Work on support escalations, resolving or routing issues based on logs and traces.
  • Maintain and manage infrastructure using Terraform and Azure DevOps.

Benefits

  • Benefits include health, dental, and vision insurance.
  • 401(k) retirement plan with company match.
  • Flexible work hours and remote work options.
  • Professional development opportunities and training.
  • Generous vacation and paid time off policy.
Full Job Description
Summary:

Delinea is looking for a Senior Site Reliability Engineer to join our Cloud Engineering team. You will own the availability and performance of several production SaaS services running in Azure and AWS, including our FedRAMP High environment. This is a hands-on role focused on measuring reliability, tuning detection, responding to incidents, and removing the manual work that keeps engineers awake at night. Prior experience in a government cloud environment is welcome but not required; we will teach you the compliance side.

What You'll Do:
  • Own reliability for a set of production SaaS services end to end: availability, performance, and capacity.
  • Define service-level indicators, set service-level objectives and error budgets, and use them to prioritize reliability work with engineering and product teams.
  • Build and tune monitoring in Datadog and Azure Monitor, including threshold, composite, and anomaly detection monitors, synthetic checks, dashboards, and alert routing that shorten time to detect and cut alert noise.
  • Automate response. Wire monitors into remediation workflows for known failure conditions and replace manual runbook steps with code.
  • Participate in an on-call rotation, lead incident response for high-severity events, and coordinate resolution across support, engineering, and product.
  • Write post-incident reviews and customer-facing root cause analyses, then drive the preventive actions to completion rather than filing and forgetting them.
  • Work support escalations: reproduce the issue, diagnose it from logs, traces, and network captures, resolve it or route it with evidence, and feed recurring patterns back into the product backlog.
  • Build and maintain infrastructure as code with Terraform and Azure DevOps pipelines.
  • Administer the web application firewall: rule tuning, rate limiting, false positive triage, and coordination with the security team.
  • Own the cost of our observability platform: Datadog index and retention spend, custom metric and APM volume, log ingestion rates, and S3 and blob storage for archived logs. Decide what gets indexed, sampled, or archived without losing the signal needed to troubleshoot.
  • Operate within our FedRAMP High environment following established change control and continuous monitoring processes, and contribute to the runbooks and standard operating procedures that support it.
  • Improve how the team runs on-call: rotation design, escalation paths, alert quality, runbook coverage, and handoff between regions.
  • Partner across Support, Security, Product, and Development to make sure new services ship with monitoring, runbooks, and SLOs in place.


What You'll Bring:
  • 8+ years in Site Reliability Engineering, DevOps, cloud operations, or production engineering for a SaaS product.
  • Hands-on Azure experience across AKS, App Service, Azure SQL, Redis, Service Bus, Front Door, and Storage, including cloud networking and cloud security fundamentals.
  • Production experience with an observability platform such as Datadog: metrics, logs, APM, dashboards, and monitor design. You should be comfortable writing queries, not just reading dashboards.
  • Demonstrated ownership of SLIs, SLOs, and error budgets for services you supported.
  • Incident response experience: you have run a bridge, made calls under pressure, and written the postmortem afterward.
  • Kubernetes in production, including ingress, deployments, resource limits, and troubleshooting failing workloads.
  • Infrastructure as code with Terraform, plus CI/CD pipeline creation and troubleshooting (Azure DevOps preferred).
  • Scripting in PowerShell and Python, and fluency with YAML and JSON.
  • Strong networking and web fundamentals: DNS, TLS and certificate chains, load balancing, reverse proxies, firewalls, and packet-level troubleshooting.
  • Knowledge of redundancy, backup, and disaster recovery strategies in cloud environments.
  • Clear written communication. Your RCAs will be read by customers and executives.
  • Willingness to participate in an on-call rotation covering weekends and emergencies.
  • Up to 10% travel.


We'd Love to See:
  • Exposure to a regulated or audited environment: FedRAMP, NIST 800-53, SOC 2, ISO 27001, PCI, or similar. Direct FedRAMP experience is a plus, not a requirement.
  • AWS experience alongside Azure, including CloudFormation and SES.
  • Hands-on experience with Jenkins pipeline creation and maintenance, SaltStack configuration management, and Consul for service discovery and distributed system coordination.
  • Advanced log analysis in the ELK stack, and/or CloudWatch Logs Insights QL.
  • Web application firewall administration (Imperva, Azure WAF, Cloudflare) and bot or rate-limiting rule tuning.
  • A track record of controlling observability tooling spend at scale: index tuning, sampling strategy, retention tiering, and archive and rehydration workflows.
  • Microsoft Entra ID and troubleshooting SAML and OIDC authentication flows.
  • Large-scale, multi-region, geo-redundant architectures and hands-on disaster recovery testing.
  • Experience with Jira Service Management, PagerDuty, or similar on-call and alerting tooling, and with public status page communication.
  • Running game days or chaos exercises, and mentoring engineers newer to SRE practice.


For this Job, Delinea is not considering candidates that need any type of US work authorization now or in the future. This includes, but is not limited to: F1-OPT, F1-CPT, H-1B, TN, L-1, J1, etc.

About Delinea

Delinea is a pioneer in securing identities through centralized authorization, making organizations more secure by seamlessly governing their interactions across the modern enterprise. Delinea allows organizations to apply context and intelligence throughout the identity lifecycle across cloud and traditional infrastructure, data, and SaaS applications to eliminate identity-related threats. With intelligent authorization for all identities, Delinea is the only platform that enables you to identify each user, assign appropriate access levels, monitor interaction across the modern enterprise, and immediately respond upon detecting any irregularities. The Delinea Platform enables your teams to accelerate adoption and be more productive by deploying in weeks, not months, and requiring 10% of the resources to manage compared to the nearest competitor.
Learn more about Delinea
Size
500 employees
Industry
Founded
2004

Similar Jobs

More Jobs at Delinea

More Information Technology Jobs

Find similar Senior Site Reliability Engineer - FedRAMP jobs: