About the rolePointClickCare builds cloud platforms that power safer, more connected care for millions of patients. You'll design and operate resilient infrastructure services across Azure, AWS, GCP, and on-prem environments, driving reliability, automation, and operational excellence for identity, compute, storage, messaging, and shared services with an SRE-first mindset.
What you'll do- Design and implement highly available infrastructure solutions for compute, storage, identity, messaging, and shared services
- Build and maintain Infrastructure as Code using Terraform or Pulumi; establish best practices and standards
- Automate operational workflows to eliminate toil: auto-remediation, self-healing systems, capacity planning
- Define and track SLIs and SLOs for critical services; manage error budgets and reliability targets
- Participate in the on-call rotation and lead incident response for complex infrastructure issues; conduct blameless post-mortems and drive systemic fixes
- Develop observability strategies: metrics, logs, distributed tracing, alerting frameworks
- Apply AI-assisted tooling to reduce toil and speed up investigation: log analysis, alert triage, runbook and post-mortem drafting, automation scaffolding
- Own reliability for one or more infrastructure domains end to end, driving multi-team initiatives with product engineering from problem definition through adoption
- Mentor intermediate SREs; review infrastructure changes; establish operational best practices
What you'll bringMust-haves- 5+ years of hands-on experience operating and designing cloud infrastructure
- Expert-level understanding of the core services of either Azure or AWS
- Working proficiency in at least one additional platform (Azure, AWS, or GCP)
- Experience designing and supporting production infrastructure that spans multiple cloud platforms
- 3+ years of production experience with Infrastructure as Code (Terraform, Pulumi, CloudFormation)
- Ability to design scalable, reusable IaC modules and enforce GitOps workflows
- Experience managing IaC across multiple cloud providers, including module design, state layout, and provider-specific resource differences
- Strong proficiency in at least one programming language (Python, Go, Bash) for production automation
- Demonstrated ability to write tested, maintainable automation and tooling
- Practical application of SRE principles in production environments
- Experience defining and managing SLIs/SLOs, error budgets, toil metrics
- Track record of improving system reliability (e.g., MTTR reduction, availability improvements)
- Demonstrated depth across key infrastructure services:
- Expert-level experience running Kubernetes and containerized workloads in production on managed Kubernetes (AKS, EKS, or equivalent), plus VM-based compute
- Practical experience operating a service mesh in production (Istio preferred; Linkerd or equivalent)
- Strong proficiency with enterprise identity and SSO: SAML, OAuth/OIDC, LDAP, and cloud IAM
- Working knowledge of an enterprise federation platform such as PingFederate, Entra ID, Okta, or ADFS
- Strong proficiency in storage solutions (object, block, and file storage; Kubernetes persistent volumes)
- Working knowledge of designing and operating infrastructure in a regulated environment (HIPAA, SOC 2, PCI, FedRAMP, or equivalent)
- Familiarity with audit evidence, access controls, encryption in transit and at rest, and data residency constraints
- Proven track record of measurably reducing operational toil through automation (e.g. ticket volume, manual runbook executions, hours reclaimed)
- Strong communication and documentation skills; demonstrated ability to influence engineering teams
Nice-to-haves- 2+ years in healthcare technology or highly regulated SaaS environments (HIPAA, SOC 2, HITRUST)
- Cloud certifications: Azure Solutions Architect Expert, AWS Solutions Architect Professional, GCP Professional Cloud Architect, or equivalent
- Working experience operating Kubernetes at scale across multiple clusters (CKA or CKAD certification a plus)
- Working knowledge of messaging and event-streaming platforms (Kafka, Azure Service Bus, Event Hubs, SQS, Pub/Sub)
- Practical experience building CI/CD pipelines and deployment automation (GitLab, GitHub Actions, ArgoCD)
- Familiarity with AI-assisted engineering and operations tooling: LLM-based coding assistants, AIOps, agentic incident investigation
- Judgment about where AI belongs in an operational workflow, including safe handling of sensitive data in prompts and reviewing generated changes before they reach production contribution to open-source SRE tools or infrastructure projects (GitHub profile, PRs merged)
Education- Bachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
- OR equivalent practical experience with a proven track record in infrastructure and SRE practices evidence of continuous learning and staying current with SRE and cloud-native trends
How we workBachelor's degree in Computer Science, Computer Engineering, Information Technology, or related technical field
- Transparent collaboration: We work in the open using OKRs, cross-functional retrospectives, and public roadmaps so everyone knows priorities and progress
- Blameless culture: We confront problems courageously through structured post-mortems and root cause analyses, focusing on systems improvement not individual blame
- Data-driven decisions: We use metrics, APM, and observability data to make evidence-based choices about reliability and performance investments
- Continuous learning: We learn from incidents through retrospectives and RCAs, sharing knowledge across teams to prevent repeat issues
- Outcome accountability: We're accountable for customer and business results, not just completing tasks, measuring success by impact
- Thoughtful experimentation: We hold strong opinions loosely, testing concepts and running small experiments before scaling solutions
- Iterative delivery: We think big but act small, using Scrum, delivery plans, and frequent milestones to ship incrementally and learn fast
- Inclusive environment: We create space to listen and learn, actively growing our Ally Community to support equity, belonging, and career growth for all
#LI-AV1
#LI-hybrid
$139,000 - $155,000 a year
At PointClickCare, base salary is one of the many components that make up our total rewards package. The CAD base salary range for this position is $139,000-$155 000 (not overtime eligible) + bonus + benefits. Compensation is assessed individually and aligned to experience, skills, and market context. The posted range reflects typical expectations for this role.
PointClickCare Benefits & Perks:Benefits starting from Day 1!
Retirement Plan Matching
Flexible Paid Time Off
Wellness Support Programs and Resources
Parental & Caregiver Leaves
Fertility & Adoption Support
Continuous Development Support Program
Employee Assistance Program
Allyship and Inclusion Communities
Employee Recognition ... and more!