Pizza Hut

Site Reliability Engineer I

Pizza Hut • $95K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's Degree preferred or equivalent technical experience
  • 1 to 3+ years in Site Reliability Engineering or related role
  • 1+ year experience with cloud platforms (AWS, Azure, GCP)
  • 1+ year experience with monitoring tools (Datadog, Prometheus, Grafana)
  • Familiarity with Linux systems and command-line utilities
  • Experience with containerization tools (Kubernetes, EKS)
  • Familiarity with scripting languages (Python, Go, Bash)

Responsibilities

  • Monitor KFC production environments using tools like Datadog
  • Participate in investigation and troubleshooting of production incidents
  • Support and improve cloud-based infrastructure
  • Contribute to infrastructure deployment using Terraform and Ansible
  • Develop scripts and automation to reduce operational toil
  • Implement reliability improvements for system performance
  • Assist with deployment readiness and production validation
  • Maintain operational documentation and contribute to post-incident analysis

Benefits

  • Comprehensive insurance coverage (medical, dental, vision)
  • Short-term and long-term disability insurance
  • Life insurance and accident coverage
  • 401(k) plan enrollment
  • 4 weeks of vacation plus paid sick leave and holidays
  • Half day Fridays year-round
  • 2 paid volunteer days each year
Full Job Description
Job Description

The Site Reliability Engineer I supports the reliability, availability, and operational health of KFC production systems and applications. This role works with Platform Engineering, Development, and Product teams to monitor and troubleshoot production environments, improve observability, contribute to automation, and support cloud-based infrastructure.
The SRE I is expected to operate within defined systems and services, contribute to reliability improvements, participate in incident response, and build increasing technical ownership over time.

Responsibilities
• Observability - Monitor KFC production environments using tools such as Datadog. Create and improve monitors, dashboards, alerts, logging, and other telemetry used to proactively identify reliability and performance issues.
• Incident Response - Participate in the investigation, troubleshooting, and restoration of production systems. Use metrics, logs, traces, and application telemetry to identify issues and support technical resolution and post-incident actions.
• Cloud & Infrastructure - Support and improve cloud-based and containerized infrastructure, including compute, networking, application services, and supporting platform components.
• Infrastructure as Code - Contribute to the deployment and maintenance of infrastructure using technologies such as Terraform, Ansible, and GitLab in partnership with SRE and Platform Engineering teams.
• Automation - Develop basic scripts, tools, and automated workflows using technologies such as Python, Go, Bash, and REST APIs to reduce repetitive operational work and toil.
• Reliability Engineering - Contribute to improvements in system availability, scalability, resilience, and performance by identifying operational gaps and implementing defined reliability improvements.
• Deployment & Production Support - Assist with deployment readiness, production validation, change verification, and monitoring following application or infrastructure releases.
• Documentation & Continuous Improvement - Maintain operational runbooks and contribute to root cause analysis, corrective actions, and improvements identified through incidents and service health reviews.
• Participate in the SRE 24/7 on-call rotation to detect, respond to, and resolve production issues.

Qualifications
• Education/Certifications - Bachelor's Degree preferred or equivalent technical experience.
• Experience
o 1 to 3+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, infrastructure, application support, software engineering, or a related technical role.
o 1+ year's experience with cloud platforms such as AWS, Azure, or GCP, including compute and networking concepts.
o 1+ year's experience with monitoring and observability platforms such as Datadog, Elastic, Dynatrace, Prometheus, or Grafana.
o 1+ year's experience with Incident and Problem Management tools such as ServiceNow, Jira, Confluence, Incident.io.
o Familiarity with Linux-based systems and command-line utilities.
o Working knowledge of networking fundamentals, including DNS, load balancing, routing, and TLS.
o Experience with containerized environments and orchestration tools such as Kubernetes or EKS.
o Familiarity with scripting or programming languages such as Python, Go, Bash, or REST APIs.
o Familiarity with Infrastructure as Code tools such as Terraform, Ansible, or CloudFormation.
o Familiarity with CI/CD tools and practices, including GitLab.
o Strong troubleshooting, analytical, communication, and collaboration skills.

Salary Range: 95,700 - 120,000

Benefits: Employees (and their eligible family members) may enroll in the following types of insurance coverage: medical, dental, vision, legal, and accidental death and dismemberment, as well as FSA/HSA (depending on enrolled medical plan). Yum! also provides short-term disability, long-term disability, and life insurance. Employees may enroll in our 401(k) plan. Yum! provides 4 weeks of vacation, paid sick leave, 10 paid holidays, a floating day off, half day Fridays year-round and 2 paid days for volunteer time each calendar year. To learn more about working at Yum! -Click here.

At Yum!, one of our core values is to Believe in ALL People. This means seeing the value in everyone and unlocking their full potential to be their best self.

Similar Jobs

More Jobs at Pizza Hut

More Information Technology Jobs

Find similar Site Reliability Engineer I jobs: