H-E-B

Senior Cloud Engineer

H-E-B$120K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in a technical field or equivalent experience.
  • 5+ years in Software Engineering, Platform Engineering, or related fields.
  • Strong experience managing large-scale cloud platforms and distributed systems.
  • Extensive experience with Databricks and large-scale data processing environments.
  • Deep expertise in AWS services including EC2, S3, and Lambda.
  • Strong programming skills in Python; SQL experience is a plus.
  • Hands-on experience with Terraform and Infrastructure as Code practices.

Responsibilities

  • Design and maintain highly available and scalable data platform infrastructure.
  • Develop and manage Infrastructure as Code solutions using Terraform.
  • Collaborate with cross-functional teams to optimize platform performance and reliability.
  • Implement comprehensive monitoring and alerting strategies aligned with business goals.
  • Troubleshoot and optimize distributed storage and compute systems in cloud environments.
  • Lead root cause analyses and develop preventative solutions for systemic risks.
  • Establish best practices for platform engineering and operational excellence.

Benefits

  • Collaborative work environment with cross-functional teams.
  • Opportunities for continuous learning and professional development.
  • Flexible working hours to support business-critical initiatives.
  • Exposure to cutting-edge technologies and best practices in platform engineering.
Full Job Description
Responsibilities

As a Senior Site Reliability Engineer supporting H-E-B's Data Platform, you will be responsible for the reliability, scalability, performance, and operational excellence of cloud-native data infrastructure and services.

Key Responsibilities
  • Design, implement, and maintain highly available, resilient, and scalable data platform infrastructure.
  • Develop and manage Infrastructure as Code (IaC) solutions using Terraform and other automation tools.
  • Partner closely with Data Engineers, Data Scientists, Analysts, and Software Engineers to optimize platform reliability and performance.
  • Develop comprehensive monitoring, alerting, observability, SLO, and capacity planning strategies aligned with business objectives.
  • Monitor, troubleshoot, and optimize distributed storage, compute, and streaming systems across cloud-based environments.
  • Lead root cause analysis efforts, identify systemic risks, and implement preventative solutions.
  • Establish and champion best practices for platform engineering, reliability engineering, security, and operational excellence.
  • Improve system resiliency through architecture reviews, fault tolerance strategies, and performance optimization initiatives.
  • Build and maintain CI/CD pipelines that support rapid, reliable deployments.
  • Implement security best practices and ensure compliance with enterprise and industry standards.
  • Drive automation initiatives that reduce operational overhead and improve platform scalability.
  • Contribute to long-term platform and reliability roadmaps.
  • Stay current with emerging technologies and recommend innovative solutions that enhance platform capabilities.
Who You Are Minimum Qualifications
  • Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related technical field (or equivalent practical experience).
  • 5+ years of experience in Software Engineering, Platform Engineering, Site Reliability Engineering, Cloud Engineering, or related disciplines.
  • Strong experience managing and supporting large-scale cloud platforms and distributed systems.
  • Extensive experience with Databricks and large-scale data processing environments.
  • Deep expertise in AWS services, including EC2, S3, VPC, IAM, Lambda, CloudFormation, and related cloud technologies.
  • Strong programming and automation experience with Python; experience with SQL is preferred.
  • Hands-on experience with Terraform and Infrastructure as Code practices.
  • Experience with distributed data and compute technologies such as Apache Spark, streaming platforms, and ETL workflows.
  • Strong understanding of software engineering principles with an emphasis on reliability, scalability, observability, and performance optimization.
  • Experience implementing monitoring, logging, observability, and incident response practices.
  • Excellent analytical, troubleshooting, and problem-solving abilities.
  • Strong communication skills with the ability to collaborate across technical and business teams.
Preferred Qualifications
  • Experience with additional cloud platforms such as Google Cloud Platform (GCP) or Microsoft Azure.
  • Experience with containerization and orchestration technologies such as Docker and Kubernetes.
  • Familiarity with CI/CD platforms and deployment automation tools.
  • Experience supporting data lake, lakehouse, or modern data platform architectures.
  • AWS, Terraform, Kubernetes, Databricks, or other relevant industry certifications.
  • Experience mentoring engineers and leading cross-functional technical initiatives.
What Makes You Successful
  • You take ownership of reliability and operational outcomes across complex distributed systems.
  • You proactively identify risks and drive long-term solutions rather than short-term fixes.
  • You can balance technical excellence with practical business needs.
  • You thrive in fast-paced environments and effectively manage competing priorities.
  • You influence engineering culture through collaboration, mentorship, and continuous improvement initiatives.
  • You are passionate about automation, scalability, and creating exceptional platform experiences for engineering teams.
Working Conditions
  • Function in a fast-paced retail and technology environment.
  • Travel by car or plane, including occasional overnight stays.
  • Sit for extended periods while working at a computer.
  • Participate in on-call rotations and support critical production systems as needed.
  • Work flexible hours when required to support business-critical initiatives.

The responsibilities listed above describe the general nature and level of work performed and are not intended to be an exhaustive list of all duties, responsibilities, or qualifications associated with this position.

About H-E-B

H-E-B is a privately held supermarket chain based in San Antonio, Texas, with more than 340 stores throughout the U.S. state of Texas, as well as in northeast Mexico. The company also operates Central Market, an upscale organic and fine foods retailer. As of 2021, the company has a total revenue of $32 billion. H-E-B was named Retailer of the Year in 2010 by Progressive Grocer.
Learn more about H-E-B
Size
120,000 employees
Industry

Similar Jobs

More Jobs at H-E-B

More Information Technology Jobs

Find similar Senior Cloud Engineer jobs: