Qualifications
Responsibilities
Benefits
THE ROLE
The SRE Principal Engineer III will work a hybrid schedule, with a requirement to be onsite at our Torrance, CA facility at least two days per week or more if needed, while also having the flexibility to work remotely. This role is responsible for leading, designing, and implementing robust Site Reliability Engineering (SRE) practices to ensure high availability, scalability, and resilience of critical business systems and applications. The SRE Principal Engineer III will focus on improving system reliability through automation, monitoring, and performance tuning, working closely with development and operations teams to champion a culture of continuous improvement and operational excellence.The SRE team consists of: 60 60 60SRE Engineers 60 60 60Deployment Automation 60 60 60Incident Response and Postmortem Analysis 60 60 60Observability and MonitoringThis role will drive the adoption of best practices in multi-cloud and hybrid-cloud platforms, managing services from major cloud providers like Microsoft Azure, Amazon AWS, Oracle OCI, Google GCP, and Alibaba Cloud. The SRE Principal Engineer III will focus on automation, incident management, performance monitoring, and optimizing infrastructure to support scalable, reliable systems. The position will also be responsible for fostering collaboration between development, operations, and security teams to streamline system operations across the organization.
HOW YOU WOULD CONTRIBUTE:
60 60 60Lead the implementation and optimization of SRE practices, ensuring system reliability, performance, and scalability. 60 60 60Architect and maintain automation for infrastructure provisioning, deployment, and incident response. 60 60 60Establish and implement SLOs (Service Level Objectives) and SLIs (Service Level Indicators) for key services. 60 60 60Collaborate with development teams to design and deliver reliable software systems, ensuring that production environments are optimized for uptime and performance. 60 60 60Create and maintain monitoring, alerting, and observability solutions to provide real-time insights into system health and performance. 60 60 60Respond to production incidents, perform root cause analysis, and implement corrective measures to prevent recurrence. 60 60 60Continuously improve system performance, capacity planning, and reliability through infrastructure tuning and automation. 60 60 60Facilitate post-incident reviews, fostering a blameless culture that focuses on learning from incidents. 60 60 60Collaborate with security teams to ensure infrastructure meets compliance, security standards, and best practices. 60 60 60Champion a collaborative environment across development, operations, and security teams to enhance operational efficiency and knowledge sharing. 60 60 60Drive the adoption of automation tools and frameworks to minimize manual intervention and optimize systems.
QualificationsSkills Required: 60 60 60Proven expertise in SRE practices, with a focus on automation, incident management, observability, and infrastructure scalability. 60 60 60Extensive knowledge of cloud platforms (Azure, AWS, GCP, Alibaba) and hybrid-cloud environments, with a focus on reliability and performance optimization. 60 60 60Experience with automation tools and scripting languages, such as Python, Go, Terraform, or Ansible, for leading infrastructure and incident response. 60 60 60Strong understanding of containerization (Docker, Kubernetes) and orchestration systems. 60 60 60Solid grasp of monitoring and observability tools (Prometheus, Grafana, Dynatrace, Splunk) to ensure real-time system health monitoring. 60 60 60Expertise in capacity planning, performance tuning, and failure management techniques. 60 60 60Strong background in incident management, root cause analysis, and postmortem processes to improve system resilience. 60 60 60Deep understanding of security and compliance requirements, and the ability to ensure production environments meet industry standards. 60 60 60Experience with Agile and DevOps methodologies to ensure fast, reliable delivery of services.
Experience Required: 60 60 6010+ years of experience in IT, with a focus on SRE, DevOps, or infrastructure engineering roles. 60 60 60Extensive hands-on experience with cloud infrastructure management and automation tools such as Terraform, CloudFormation, or equivalent. 60 60 60Proficiency in scripting and automation languages like Python, Bash, Go, or Ruby for infrastructure automation. 60 60 60Proven experience in managing large-scale systems, ensuring reliability, high availability, and scalability. 60 60 60Expertise in container orchestration technologies, including Kubernetes, OpenShift, and Docker Swarm. 60 60 60Deep knowledge of monitoring and observability platforms (Prometheus, Grafana, ELK, Dynatrace), including experience building and maintaining alerting and dashboard systems. 60 60 60Strong understanding of version control systems and CI/CD practices to optimize code deployment as it relates to infrastructure. 60 60 60Demonstrated ability to optimize performance in multi-cloud and hybrid-cloud environments, ensuring uptime and performance at scale.
Education Required: 60 60 60Bachelor99s degree in computer science, Information Technology, or related field, or equivalent experience.
Certificates / Training Preferred: 60 60 60Relevant cloud certifications such as AWS Certified Solutions Architect, Azure Solutions Architect Expert, or Google Cloud Professional Cloud Architect. 60 60 60SRE-related certifications like Certified Kubernetes Administrator (CKA) or Google Professional Cloud DevOps Engineer.
About Herbalife
Similar Jobs

More Jobs at Herbalife




More Information Technology Jobs

