Cloudbeds

Site Reliability Engineer

Cloudbeds • $120K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience as a DevOps or SRE in the AWS environment.
  • 5+ years of experience managing Kubernetes (EKS) and Helm charts.
  • Proficient in CI/CD pipeline development using ArgoCD and GitHub actions.
  • Strong skills in infrastructure-as-code with Terraform.
  • Experience in Observability and Monitoring using tools like Grafana, Prometheus, DataDog, and Cloudwatch.
  • Familiarity with Incident Management and Root Cause Analysis (RCA).
  • Solid understanding of networking principles and configurations (VPC, Security Groups).

Responsibilities

  • Architect and implement scalable AWS cloud solutions.
  • Maintain and support Kubernetes (EKS) clusters and related infrastructure.
  • Enhance the CICD process using ArgoCD and GitOps methodologies.
  • Automate platform deployments with Terraform methodologies.
  • Develop and improve monitoring and observability systems for the platform.
  • Engage in Incident Management and conduct efficient Root Cause Analysis.
  • Optimize system performance and troubleshoot issues quickly.

Benefits

  • Collaborative team culture that fosters innovation and autonomy.
  • Opportunity to influence architectural decisions in a cutting-edge tech environment.
  • Work remotely with flexible time management in a global team.
  • Engage in continuous improvement and automation initiatives.
Full Job Description
As a Site Reliability Engineer, you'll be the guardian of our platform's reliability and performance, ensuring millions of hospitality transactions flow seamlessly across the globe. You'll architect and implement scalable AWS cloud solutions that keep the most ambitious hotels running 24/7, while fostering a culture of automation, resilience, and continuous improvement across our engineering teams.

Our SRE Team:

We're a bottom-up, collaborative team that thrives on healthy debate and shared ownership of our infrastructure. You'll have endless opportunities to influence architecture decisions while working with cutting-edge cloud technologies at scale. We believe the best solutions come from engineers who are empowered to innovate, experiment, and challenge the status quo.

What You Bring to the Team:
  • Design and implement a reliable and scalable AWS architecture to meet the needs of the organization.
  • Maintain and support highly loaded Kubernetes (EKS) clusters and infrastructure-related components.
  • Support the CICD process with ArgoCD and GitOps.
  • Automate the platform deployments with Terraform infrastructure-as-code.
  • Develop and continuously improve product Observability and Monitoring systems based on the Grafana, Prometheus, DataDog, and Cloudwatch.
  • Respond and participate with Incident Management and Root Cause Analysis, ensuring minimal impact on services.
  • Optimize system performance and troubleshoot issues as they arise.
  • Collaborate with development teams to establish monitoring best practices and ensure systems meet reliability targets.
  • Collaborate with security teams to implement and maintain security best practices.
  • Infrastructure support rotation providing guidance to other engineering teams.

What Sets You Up for Success:
  • 5+ years of experience as a DevOps or SRE working within the AWS ecosystem.
  • 5+ years of experience with Kubernetes (EKS) and Helm charts.
  • Experience with designing, building, and supporting CI/CD pipelines with ArgoCD and GitHub actions.
  • Experience with infrastructure-as-code methodologies with Terraform.
  • Experience with Observability and Monitoring with Grafana, Prometheus, DataDog, and Cloudwatch.
  • Experience with Incident Management, full stack troubleshooting, performance analysis and root cause analysis (RCA).
  • Experience with Web application systems such as Nginx, Ingress controllers, load balancing and Content Delivery Networks.
  • Experience with Databases (MySQL, PostgreSQL, Aurora) and Middleware technologies (Redis, Memcached and SQS)
  • Good networking skills with VPC, Security Groups and Network ACLs.
  • Ability to work remotely and manage your own time in a global team.
  • Good written and verbal communication in English.
  • Bachelor's degree in Computer Science or equivalent experience.

Bonus Skills to Stand Out:
  • Advanced experience with Database Administration (Aurora, MySQL, PostgreSQL).
  • Experience working in a PCI-compliant environment.
  • Experience working with Kong API Gateway.

Compensation: Depending on your skills and experience, you can expect your annual compensation to be between $120,000 - $150,000

#LI-REMOTE #LI-IK1

Work Authorization: Please note that applicants must be currently authorized to work in the location where the position is located without requiring visa sponsorship. At this time, Cloudbeds is unable to provide sponsorship for work visas

About Cloudbeds

Cloudbeds is a hospitality management software company that provides a suite of tools to help hotels, hostels, and vacation rentals manage their operations. The company's platform includes features such as property management, channel management, and booking engine, as well as integrations with other hospitality software providers. Cloudbeds was founded in 2012 by Adam Harris and Richard Castle.
Learn more about Cloudbeds
Size
500 employees
Industry
Founded
2012

Similar Jobs

More Jobs at Cloudbeds

  • Cloudbeds
    Site Reliability Engineer
    $120K — $150K *
    Virginia, MN 55792 (Saint Louis County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: