Lead Site Reliability Engineer

Lumen

$105K — $155K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8 years in DevOps/Platform/SRE roles; 2 years leading teams or cross-functional initiatives.
  • Deep expertise in Azure (Public/Gov) cloud services and resource management.
  • Hands-on with AKS, Helm 3, Flux CD (GitOps), Docker.
  • Experience with Prometheus & Grafana (metrics, alerting, dashboards).
  • Proficiency in Nginx and Kong (ingress/API gateway).
  • CI/CD automation using Jenkins and GitHub Actions.
  • Strong development skills for automation & debugging (Python, Go, or TypeScript; Bash/PowerShell).
  • Proven experience with IaC using Terraform (or similar tools like Bicep), policy-as-code, and secure SDLC integrations.
  • Ability to obtain GSA Tier 2 Suitability Clearance.

Responsibilities

  • Design, build, and operate AKS clusters with enterprise guardrails.
  • Implement GitOps with Flux CD; define Helm chart standards.
  • Own container platform best practices including vulnerability scanning and base image governance.
  • Configure Nginx and Kong including routing and WAF integration.
  • Stand up Prometheus for metrics scraping and alerting.
  • Lead Jenkins pipeline automation and GitHub Actions workflows.
  • Build infrastructure-as-code and policy-as-code pipelines using Terraform.
  • Architect secure integration with Azure Key Vault for secret management.
  • Design and implement Azure Networking components.

Benefits

  • Broad range of Health, Life, and Voluntary Lifestyle benefits.
  • Perks to enhance physical, mental, emotional, and financial wellbeing.
Full Job Description
The Role

Lumen is seeking a Lead Site Reliability Engineer (SRE) who will be a catalyst for transformational change and operational excellence. The ideal candidate not only builds and operates systems but proactively identifies gaps, challenges legacy approaches, and delivers high-impact changes that improve customer experience. You will own platform engineering standards across Azure Public and Azure Government clouds, lead the design of CI/CD and observability systems, and partner closely with development teams to deliver scalable, secure, and compliant high-availability services.

Location

This role is designated as a fully remote position within the United States.

The Main Responsibilities

  • Design, build, and operate AKS clusters with enterprise guardrails (RBAC, pod security policies, node pools, autoscaling, upgrade strategies).
  • Implement GitOps with Flux CD; define Helm chart standards and manage lifecycle via Helm 3.
  • Own container platform best practices: Docker image standards, multi-stage builds, SBOM, vulnerability scanning, and base image governance.
  • Configure Nginx (ingress controller) and Kong (API gateway) including routing, rate limiting, mutual TLS, JWT/OAuth2 auth, and WAF integration.
  • Stand up Prometheus for metrics scraping, alert rules, and exporters; integrate Grafana for dashboards and SLO/SLA visualization.
  • Lead Jenkins pipeline automation and GitHub Actions workflows for build/test/deploy, with reusable templates and environment promotion strategies.
  • Build infrastructure-as-code and policy-as-code pipelines using Terraform (or similar tools like Bicep), with automated validations, PR gates, and compliance checks.
  • Architect and enforce secure integration with Azure Key Vault for secret, key, and certificate management; automate cert-manager/issuer (ACME/PFX).
  • Design and implement Azure Networking: VNets, NSGs, route tables, Private Link, Firewall rules, egress control, and traffic segregation in hub-spoke models.
  • Operate and integrate Azure Blob Storage, Azure Redis Cache, PostgreSQL, Azure Cosmos DB (Mongo API), and Azure Service Bus.


What We Look For in a Candidate

Required Qualifications:
  • 8 years in DevOps/Platform/SRE roles; 2 years leading teams or cross-functional initiatives.
  • Deep expertise in Azure (Public/Gov) cloud services and resource management.
  • Hands-on with AKS, Helm 3, Flux CD (GitOps), Docker.
  • Experience with Prometheus & Grafana (metrics, alerting, dashboards).
  • Proficiency in Nginx and Kong (ingress/API gateway).
  • CI/CD automation using Jenkins and GitHub Actions.
  • Strong development skills for automation & debugging (Python, Go, or TypeScript; Bash/PowerShell).
  • Proven experience with IaC using Terraform (or similar tools like Bicep), policy-as-code, and secure SDLC integrations.
  • Ability to obtain GSA Tier 2 Suitability Clearance.
Preferred Qualifications:
  • Experience in Azure Government with FedRAMP controls and ATO support.
  • Certifications: AZ-104, AZ-305, AZ-400, CKA/CKAD, Terraform Associate.
Clearance Requirement:

Must be willing to go through GSA Tier 2 Suitability Clearance process post onboarding.

Compensation

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY
$111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI
$116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process. Learn more about Lumen's:Benefits
Bonus Structure

#LI-Remote

#LI-VK1

Requisition #: 343285

Similar Jobs

More Jobs at Lumen

More Information Technology Jobs

Find similar Lead Site Reliability Engineer jobs: