Frontier Airlines

Lead Engineer - Site Reliability

Frontier Airlines$110K — $146K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science, Engineering, Information Technology, or a related discipline.
  • 10+ years in software engineering, cloud architecture, or infrastructure engineering.
  • 5+ years of experience with AWS architecture and cloud transformations.
  • Demonstrated leadership in large-scale cloud migrations and modernization.
  • Expertise in Kubernetes, microservices, and distributed systems.
  • Proficient in DevSecOps, CI/CD pipelines, and Infrastructure as Code.
  • Strong understanding of Site Reliability Engineering practices.

Responsibilities

  • Lead the design of resilient cloud platforms using AWS and cloud-native technologies.
  • Define reliability engineering strategies, SLOs, and SLIs for critical services.
  • Architect Kubernetes platforms for containerized applications.
  • Establish reliability standards across the organization.
  • Collaborate with various teams to enhance service reliability and lower operational risks.
  • Drive DevSecOps adoption and Infrastructure as Code practices.
  • Lead incident responses and continuous improvement initiatives.

Benefits

  • Flexible work arrangements to promote work-life balance.
  • Training and development programs for skill enhancement.
  • Opportunities for professional certifications.
  • Supportive company culture prioritizing collaboration and innovation.
Full Job Description
What Will You Be Doing?

As a Lead Site Reliability Engineer (SRE), you will serve as a senior technical leader responsible for ensuring the reliability, scalability, performance, and operational excellence of mission-critical technology platforms. You will partner with engineering, platform, security, and operations teams to build resilient systems, drive automation, improve service availability, and establish reliability-focused engineering practices across the organization.

This role requires deep expertise in AWS, Kubernetes, observability, distributed systems, DevSecOps, and Site Reliability Engineering principles, with the ability to influence engineering teams and operational strategies across the enterprise.

Essential Functions

  • Lead the design and implementation of highly available, resilient, and scalable cloud platforms leveraging AWS services and cloud-native technologies.
  • Define and execute reliability engineering strategies, service level objectives (SLOs), service level indicators (SLIs), and error budgets across critical business services.
  • Architect and operate Kubernetes-based platforms supporting containerized and microservices-based applications.
  • Establish enterprise standards for reliability, availability, disaster recovery, resiliency testing, and operational excellence.
  • Partner with software engineering, platform engineering, security, and operations teams to improve service reliability and reduce operational risk.
  • Drive adoption of DevSecOps practices, CI/CD automation, Infrastructure as Code (IaC), and self-service platform capabilities.
  • Define and implement observability standards including monitoring, logging, distributed tracing, synthetic monitoring, and operational intelligence.
  • Lead incident response, post-incident reviews, root cause analyses (RCA), and continuous improvement initiatives.
  • Identify and eliminate single points of failure through proactive reliability engineering and resilient architecture practices.
  • Design and implement automated remediation, self-healing capabilities, and proactive alerting solutions to improve operational efficiency.
  • Lead capacity planning, performance engineering, availability management, and scalability assessments.
  • Partner with FinOps and engineering teams to optimize cloud resource utilization while maintaining reliability objectives.
  • Evaluate emerging technologies, AIOps, Generative AI, and intelligent automation solutions to improve platform reliability and operational effectiveness.
  • Mentor SREs, platform engineers, and software engineers in reliability practices, observability, automation, and operational excellence.
  • Develop executive-level reliability roadmaps, operational strategies, and platform investment recommendations.


Additional Responsibilities

  • Act as a trusted advisor to executive leadership on reliability strategy, operational risk management, and service resiliency initiatives.
  • Support platform evaluations, architecture reviews, and technology decisions through the lens of reliability, scalability, and maintainability.
  • Provide technical leadership during major incidents, business-critical outages, and high-severity escalations.
  • Assist with compliance initiatives including PCI-DSS, SOC 2, security governance, and operational resilience controls.
  • Contribute to enterprise-wide cloud transformation, platform engineering, and application modernization programs.
  • Champion a culture of reliability, automation, operational ownership, and continuous improvement across engineering teams.


Qualifications

  • Bachelor's degree in computer science, Engineering, Information Technology, or a related discipline.
  • 10+ years of experience in software engineering, cloud architecture, infrastructure engineering, or enterprise architecture.
  • 5+ years of hands-on AWS architecture and cloud transformation experience.
  • Proven success leading large-scale cloud migrations and modernization initiatives.
  • Experience designing and supporting highly available, mission-critical, customer-facing platforms.
  • Deep understanding of Kubernetes, container orchestration, microservices, and distributed systems.
  • Extensive experience with DevSecOps, CI/CD pipelines, Infrastructure as Code, and automation frameworks.
  • Strong knowledge of Site Reliability Engineering (SRE), operational excellence, and platform reliability practices.
  • Experience implementing cloud governance, FinOps, and cost optimization programs.


Preferred Qualifications

  • Experience supporting large-scale enterprise, eCommerce, aviation, travel, SaaS, or high-volume digital platforms.
  • Experience building internal developer platforms and platform engineering capabilities.
  • AWS Professional and/or Kubernetes certifications.
  • Experience with AIOps, intelligent automation, and reliability analytics.


Knowledge, Skills and Abilities

  • Reliability & Platform Technologies

    • Amazon Web Services (AWS)
    • Kubernetes (EKS)
    • Container Platforms
    • Platform Engineering
    • Infrastructure as Code (Terraform, CloudFormation)
    • Cloud Networking and Security
    • High Availability & Disaster Recovery
    • Performance Engineering & Capacity Planning

  • Site Reliability Engineering

    • Service Level Indicators (SLIs)
    • Service Level Objectives (SLOs)
    • Error Budget Management
    • Reliability Engineering
    • Availability & Resiliency Design
    • Chaos Engineering
    • Incident Response & Disaster Recovery
    • Root Cause Analysis (RCA)


  • Observability & Operations

    • Monitoring & Alerting
    • Distributed Tracing
    • Centralized Logging
    • Synthetic Monitoring
    • Operational Intelligence
    • AIOps & Event Correlation
    • Incident & Problem Management
    • Change & Release Management

  • DevSecOps & Automation

    • CI/CD Automation
    • DevSecOps Practices
    • Infrastructure Automation
    • Platform Automation
    • Self-Healing Systems
    • Configuration Management
    • Reliability Automation

  • Leadership & Business Skills

    • Operational Strategy & Reliability Roadmaps
    • Cloud Financial Management (FinOps)
    • Executive Communication
    • Risk Management & Mitigation
    • Technical Leadership Without Direct Authority
    • Cross-Functional Collaboration
    • Major Incident Leadership
    • Continuous Improvement Leadership


Equipment Operated

Standard office equipment, including PC, copier, fax machine, printer

Work Environment

Typical office environment, adequately heated and cooled

Physical Effort

Generally, not required.

Supervision Received

General Direction: The incumbent normally receives little instruction on day-to-day work and receives general instructions on new assignments.

Salary Range: $110,114 - $146,157 - Please note: this role will close on or before 10/31/26.

Positions Supervised

None

Workplace Policies

Disclaimer: The above statements are intended only to describe the general nature and level of work required of the referenced position; they are not intended to be an exhaustive list of all responsibilities, duties, and skills required of individuals in this position. Please be advised that duties and expectations of this position may be subject to change.

About Frontier Airlines

Frontier Airlines is a low-cost airline that operates flights to over 100 destinations in the United States, Mexico, and the Caribbean. The company was founded in 1994 and is headquartered in Denver, Colorado. Frontier Airlines is known for its low fares and customer-friendly policies, such as allowing passengers to bring one personal item and one carry-on bag for free. The airline has a fleet of over 100 aircraft and is constantly expanding its route network.
Learn more about Frontier Airlines
Size
4,000 employees
Industry
Founded
1994

Similar Jobs

More Jobs at Frontier Airlines

More Information Technology Jobs

Find similar Lead Engineer - Site Reliability jobs: