AlayaCare

Senior Site Reliability Specialist (SRE)

AlayaCare$90K — $120K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or advanced degree in computer science, engineering, or related fields
  • 5+ years of hands-on experience in site reliability engineering or similar roles
  • Solid experience with AWS in multi-account, multi-region configurations
  • Proficiency in Terraform and Infrastructure as Code methodologies
  • Hands-on experience with Docker and Kubernetes in production
  • Strong skills in Linux systems administration and troubleshooting
  • Familiarity with observability platforms and event-driven alerting mechanisms

Responsibilities

  • Design and maintain AWS cloud infrastructure and Kubernetes services
  • Implement infrastructure as code for reliable environments
  • Monitor production systems, addressing incidents and improving observability
  • Participate in on-call rotations, incident handling, and reviews
  • Collaborate with Product and Engineering teams to translate requirements into infrastructure solutions
  • Identify risks and recommend trade-offs for operational efficiency
  • Contribute to operational enhancements through process improvements and performance tuning

Benefits

  • Hybrid work model with 2 in-office days per week
  • Emphasis on team connection, innovation, and collaboration
  • Opportunity to work on AI-driven tools
  • Exposure to advanced SRE practices
  • Support for continuous learning and staying current with industry trends
Full Job Description
About the Role

We are seeking a Senior Site Reliability Specialist to join our SRE team. Reporting to the Engineering Manager, you is responsible for scaling AWS cloud infrastructure, evolving Kubernetes deployment pipelines, improving monitoring, alerting, and resiliency, and developing tooling that enables product teams to deliver safely and efficiently. For acquired Azure-based products the focus is on monitoring, alert triage, and runbook-driven incident response rather than greenfield platform design.

This role owns shared platform services across cloud regions, including databases, messaging, logging, search, and tenant provisioning. The Senior SRE is expected to lead major infrastructure initiatives and proof-of-concept efforts, contribute to technical planning and prioritization, as well as partners with the Product teams to reduce operational incidents.

The SRE team also develops and operates AI-driven tools to streamline runbooks, accelerate incident response, and generate operational insights from platform telemetry to improve reliability and reduce manual effort.
What You'll Do
Development, Automation, and Tooling
  • Design, build, and maintain infrastructure and platform services, including Kubernetes and observability tooling.
  • Implement infrastructure as code, configuration management, and automated testing to ensure reliable, repeatable environments.
  • Contribute to code and configuration reviews to improve scalability,maintainability, and reuse.
Reliability and Operations
  • Monitor production systems, troubleshoot issues, and improve logging, monitoring, alerting, and runbooks.
  • Participate in on-call and help desk rotations, incident response, and post-incident reviews to improve long-term reliability.
Requirements and Collaboration
  • Partner with Product, Engineering, and development teams to translate requirements into reliable and operable infrastructure solutions.
  • Identify risks across operability, security, performance, and cost, and recommend practical trade-offs.
Continuous Improvement
  • Contribute to operational quality through runbooks, security hardening, performance tuning, and process improvements.
  • Stay current with emerging SRE practices, including AI-assisted operations and modern AWS platform patterns.
What You Bring to the Team
  • Bachelor's or advanced degree in computer science, computer engineering, or related practical fields with demonstrated experience.
  • 5+ years of hands-on experience
  • Solid hands-on experience with AWS in a multi-account, multi-region environment: EKS, AWS Organizations, IAM, and KMS.
  • Strong proficiency with Terraform and Infrastructure as Code workflows, including Atlantis/GitOps, state management, and module/provider upgrades.
  • Practical experience running workloads on Docker and Kubernetes in production, including Gateway API ingress patterns and cluster lifecycle management (upgrades, addons, node provisioning).
  • Strong experience with Linux systems administration and production troubleshooting.
  • Proficiency in at least one development or scripting language such as Python, Go, or Bash.
  • Experience with an observability platform (e.g. New Relic, OpenSearch, CloudWatch, OpenTelemetry) and event-driven alerting (e.g. EventBridge, SNS, PagerDuty).
  • Knowledge of system and network security fundamentals, including WAF, least-privilege IAM, secrets management, and backup/disaster recovery.
  • Experience participating in incident management (on-call, triage, remediation, post-incident review) and writing operational runbooks.
  • Hands-on experience operating Aurora MySQL and PostgreSQL in production, including migrations, performance tuning, and backup/restore.
  • Strong communication and collaboration skills, with the ability to work effectively across technical and non-technical teams in a distributed environment.
  • Experience with cloud cost optimization (rightsizing, reserved capacity, and cost allocation tagging) is an asset.
  • Experience with Flux/ArgoCD, Karpenter, or Ray/Anyscale GPUinfrastructure is an asset.
  • Experience monitoring production systems on Azure is an asset.
  • Relevant cloud or Kubernetes certifications are an asset.
  • Bilingual in French and English is an asset
Location and Work Model

This role is based in Montreal. At AlayaCare, our hybrid model includes 2 set in-office collaboration days/week, and it is expected that team members are present in the office on those days to foster connection, innovation, and teamwork.

About AlayaCare

AlayaCare is a Canadian software company that provides cloud-based home healthcare software. The company's platform includes features such as electronic health records, scheduling, billing, and reporting, and is used by home care agencies, caregivers, and patients. AlayaCare's mission is to improve the quality of life for patients and caregivers by providing innovative technology solutions. The company has received several awards for its technology and has been recognized as one of Canada's fastest-growing companies.
Learn more about AlayaCare
Size
200 employees
Industry
Founded
2014

Similar Jobs

More Jobs at AlayaCare

More Information Technology Jobs

Find similar Senior Site Reliability Specialist (SRE) jobs: