Lead Site Reliability Engineer

Lumen

$105K — $155K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree or equivalent in engineering, computer science, or related field.
  • 8+ years in software development, systems engineering, and/or networking.
  • 5+ years of related experience required.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP).
  • Strong automation and infrastructure-as-code skills: Terraform, Ansible, and Python.
  • Experience running containerized workloads on Kubernetes.
  • Working knowledge of modern observability and monitoring tooling.

Responsibilities

  • Serve as subject matter expert for network automation platform applications and services.
  • Build and maintain the observability stack, including dashboards and visualizations.
  • Define and tune proactive alerting for system issues.
  • Lead incident response and participate in on-call rotation for service outages.
  • Drive blameless postmortems and prevent recurrence through process improvements.
  • Automate deployment pipelines and cloud infrastructure provisioning.
  • Collaborate with cross-functional teams to enhance NaaS applications.

Benefits

  • Comprehensive health benefits for physical, mental, and emotional wellbeing.
  • Life and voluntary lifestyle benefits.
  • Flexible working arrangements with a fully remote position.
  • Opportunities for mentorship and professional growth within the team.
  • Bonus structure including short-term and long-term incentives.
Full Job Description
The Role

Lumen's Network as a Service (NaaS) platform delivers on-demand networking at scale. As Lead SRE, you'll own the reliability of that platform - partnering with operations teams and development counterparts to drive technical direction and resolve systemic issues across a broad range of network topologies and applications.

You'll be accountable for platform observability, incident management, and automation, and you'll coordinate across architecture, engineering, and systems development organizations to measurably improve reliability. You'll also use AI and agentic tooling to build utilities that accelerate deployment automation, platform administration, and incident investigation.

Success in this role draws on networking fundamentals, cloud platforms, software development and troubleshooting methodology, and a bias toward automating what you'd otherwise do twice. We're looking for a change maker - someone who sees where the platform should go next and drives meaningful impact for the customers who rely on it

Location

This role is designated as a fully remote position within the United States.

The Main Responsibilities

  • Reliability & Observability
    • Serve as subject matter expert for network automation platform applications, services, and hosting environments
    • Build and maintain the observability stack: instrument services, collect and curate metrics, and create dashboards and visualizations that make system health obvious at a glance
    • Define and tune proactive alerting so issues surface before customers feel them
    • Champion core SRE principles - SLIs, SLOs, and error budgets - and advocate for resilient, fault tolerant architecture
  • Incident Management
    • Participate in an on-call rotation and lead incident response for service outages and unplanned downtime
    • Drive blameless postmortems and root cause analysis; own follow-up actions through to completion
    • Prevent recurrence through process improvements, tooling, and knowledge sharing across teams
  • Automation & Infrastructure
    • Automate deployment pipelines (CI/CD) and cloud infrastructure provisioning, scaling, and configuration using infrastructure as code
    • Develop tools and utilities that reduce toil and empower operations and development teams to manage services independently
    • Apply AI-assisted and agentic workflows to development, support, and investigation work
  • Collaboration & Leadership
    • Collaborate with cross-functional development teams to support, enhance, and scale NaaS applications
    • Provide guidance and mentorship to junior engineers
    • Maintain clear documentation for processes and architecture


What We Look For in a Candidate

Required Qualifications:
  • Bachelor's degree or equivalent in engineering, computer science, or related field.
  • 8+ years in software development, systems engineering, and/or networking
  • 5+ years of related experience required.
  • Hands-on experience with at least one major cloud platform (AWS, Azure, or GCP), including compute, networking, and identity services
  • Strong automation and infrastructure-as-code skills: Terraform, Ansible, and Python
  • Experience running containerized workloads on Kubernetes
  • Working knowledge of modern observability and monitoring tooling (e.g., Datadog, CloudWatch, Grafana, Prometheus), including building dashboards and defining alerts
  • Demonstrated experience with incident management and blameless postmortems
  • Comfort using AI-assisted development and agentic tools as part of daily engineering practice
  • Understanding of network technologies including Internet, Ethernet, IPVPN, Edge Compute, and Optical transport
  • Strong listening and communication skills; able to operate with autonomy while knowing when to escalate


Preferred Qualifications:
  • Multi-cloud experience across AWS, Azure, and GCP
  • Asynchronous programming concepts and distributed systems design
  • Zero-downtime deployment strategies
  • High availability and multi-region architectures
  • Source control and CI/CD practices at scale
  • Experience applying agentic workflows to operational support and investigation


Compensation

This information reflects the anticipated base salary range for this position based on current national data. Minimums and maximums may vary based on location. Individual pay is based on skills, experience and other relevant factors.

Location Based Pay Ranges

$105,786 - $141,047 in these states: AL AR AZ FL GA IA ID IN KS KY LA ME MO MS MT ND NE NM OH OK PA SC SD TN UT VT WI WV WY
$111,074 - $148,099 in these states: CO HI MI MN NC NH NV OR RI
$116,364 - $155,152 in these states: AK CA CT DC DE IL MA MD NJ NY TX VA WA

Lumen offers a comprehensive package featuring a broad range of Health, Life, Voluntary Lifestyle benefits and other perks that enhance your physical, mental, emotional and financial wellbeing. We're able to answer any additional questions you may have about our bonus structure (short-term incentives, long-term incentives and/or sales compensation) as you move through the selection process. Learn more about Lumen's:Benefits
Bonus Structure

#LI-Remote

#LI-VK1

Requisition #: 343429

Similar Jobs

More Jobs at Lumen

More Information Technology Jobs

Find similar Lead Site Reliability Engineer jobs: