Must Have Technical/Functional Skills:
• 7+ years of experience designing, deploying, and operating mid to large-scale enterprise or cloud environments.
• Familiarity with private cloud technologies including VMware, bare metal environments, file and block storage, as well as cloud infrastructure providers like AWS, GCP.
• Operational experience with container orchestration platforms (Kubernetes) at production scale, workload deployment via Helm, ArgoCD
• Familiarity with automation frameworks for OS image creation and deployment.
• Strong scripting and coding skills in languages such as Python, Bash, Ruby, or Scala.
• Expertise with infrastructure/configuration management tools such as Ansible, or Terraform.
• Deep understanding of systems and application design with operational trade-offs.
• Proficiency with modern Unix/Linux operating systems and distributions.
• Strong collaboration, mentoring, documentation and communication skills.
• Ability to prioritize tasks, work independently, and call out issues effectively.
• BS or MS degree in Computer Science, Engineering, or equivalent experience.
Roles & Responsibilities:
• Design, implement, and manage scalable, highly available private cloud infrastructure and orchestration platforms supporting compute and storage services.
• Automate infrastructure provisioning, configuration, and deployment using Infrastructure as Code (IaC) tools.
• Monitor, troubleshoot, and optimize platform performance, reliability, and security.
• Develop and maintain standardized OS images and deployment frameworks for bare metal, virtual machines, and cloud instances.
• Collaborate with other engineering teams to ensure platform uptime, cultivate sound engineering principles and represent our engineering values.
• Build end-to-end documentation, instrumentation, and automation to enable self-healing and resiliency.
• Participate in on-call rotations to support production environments.
• Lead and contribute to large-scale projects, promoting engineering standards and team values such as inclusivity, customer focus, and continuous improvement.
Nice to have skills:
As a Site Reliability Engineer within the Private Cloud team, you will architect, build, and support highly available, scalable, and resilient private cloud infrastructure. You will focus on automating operations, improving platform reliability, and enabling seamless service deployment for internal customers. This role requires collaboration with cross-functional teams to ensure platform stability, security, and continuous improvement aligned with mission to provide reliable compute and storage services
Salary Range: $64,000 - $130,000 a year
#LI-CM2