Lead, Site Reliability Engineer

Etraveli Group

$140K — $180K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in production infrastructure, minimum 2 in a lead role
  • Proven experience leading SRE, platform, or infrastructure teams
  • Deep production Kubernetes experience, including cluster lifecycle and networking
  • Strong GCP experience, ideally with region migrations and production traffic
  • Expertise in Infrastructure as Code and GitOps; direct experience migrating from legacy config management
  • Familiarity with observability and reliability practices like SLOs and error budgets
  • Security-oriented with experience in least-privilege access design and secrets management

Responsibilities

  • Lead the Toronto SRE team, focusing on technical guidance and career development
  • Manage incident response for shared infrastructure with a sustainable on-call rotation
  • Communicate migration risks and reliability trade-offs to senior leadership
  • Oversee GCP to EU migration strategy and execution with zero downtime target
  • Define workload placement strategy for OpenStack/Talos environment
  • Establish a consistent deployment model across platforms using GitOps
  • Coordinate cross-Atlantic network architecture and decommission legacy infrastructure

Benefits

  • Remote work flexibility and collaboration across multiple time zones
  • Access to industry-standard tools and cutting-edge technologies
  • Opportunity to lead an innovative infrastructure transition
  • Support for professional development and mentoring of SRE engineers
  • Participation in building a robust, secure, and compliant environment for critical systems
Full Job Description
The role

Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on. This spans three Kubernetes environments today - GKE on GCP as the primary platform for booking and search workloads, a multi-tenant Kubernetes-on-OpenStack cluster running Talos Linux for content acquisition and SRE tooling, and an isolated PCI-compliant environment for payment processing - plus a substantial VM footprint, self-hosted Concourse CI, and a Prometheus, Thanos, and Grafana observability stack.

We are in the middle of a major infrastructure transition. Our Toronto data centre closes in September 2026, and the first phase of migrating the Tripstack Platform to a new OpenStack and Talos-based Kubernetes environment in Gothenburg, Sweden - built in partnership with our parent company Etraveli Group under a shared responsibility model - is in progress. Over the next year we intend to further optimize workloads across our environments. The most immediate priority for this role is moving our GCP footprint from the US to the EU by the end of 2026 with a zero downtime objective.

This role has two parts. First, lead the Toronto SRE team day-to-day: on-call, incident management, delivery, and developing engineers who can communicate risk and progress clearly to leadership. Second, shape the overall strategy for application deployments across GCP and Gothenburg - which workloads move, which stay, where the PII boundary sits, and the deployment patterns and reliability standards that apply across both.

This is a hands-on leadership role on a distributed team spanning Toronto, Pune, and Kraków, working closely with engineering counterparts in Stockholm and Gothenburg.
Responsibilities

Lead the Toronto SRE team
  • Provide day-to-day technical and people leadership for SRE engineers based in Canada: priorities, delivery, code and change review, career development
  • Own on-call and incident management for shared infrastructure: sustainable rotations, current runbooks, and post-incident reviews that result in concrete improvements
  • Develop engineers who communicate well with leadership: your team should be able to present migration risks and reliability trade-offs to senior stakeholders directly
  • Coordinate closely with SRE and engineering colleagues in Pune and Kraków so that ownership and handoffs across time zones are clearly defined

Shape deployment strategy across GCP and Gothenburg
  • Own the GCP US-to-EU region migration end to end: planning, sequencing, cutover, and validation for live production traffic, with completion targeted by the end of 2026
  • Define the workload placement strategy: which workloads move to the OpenStack/Talos environment in Gothenburg, which remain on GCP, and how the requirement that PII stays in GCP EU is enforced
  • Establish a consistent deployment model across both platforms: Concourse pipelines, Helm, and GitOps patterns that work the same way whether the target is GCP or Gothenburg
  • Consolidate the two identity systems we operate today - LDAP/Keystone for the OpenStack estate and GCP IAM for cloud - into a clear and consistent access model
  • Define the standard for how application teams onboard workloads: node pools, namespaces, quotas, network policy, and secrets management, all documented and consistent

Deliver the data centre migration
  • Co-own execution of the Gothenburg migration with our Sr. Manager, SRE and ETG ITOPS counterparts, within the shared responsibility model - Tripstack owning the Kubernetes control plane and everything above the hypervisor, adopting ETG standards below it
  • Complete phase one, then plan and execute the subsequent migration waves - moving business-critical services from single-homed to fully redundant, with failover scenarios enabled and tested regularly
  • Own the network architecture of a cross-Atlantic platform: peering between Canada, Sweden, and GCP EU, and latency requirements for critical paths
  • Decommission legacy infrastructure as part of the migration: legacy Terraform, Puppet-managed VMs, and VM-based tooling that has a Kubernetes-native replacement

Raise the reliability bar
  • Build a formal SLO framework: SLIs, error budgets, and dashboards for our critical APIs, so that reliability decisions are based on data
  • Extend the Prometheus / Thanos / Grafana / alerting stack so that both platforms, and the migration itself, are fully observable
  • Make post-incident follow-through, capacity planning, and change safety standard practice across the engineering organization

Own the security posture of the infrastructure
  • Partner with our Security team on infrastructure security: least-privilege access across both identity systems, secrets management through Vault, network segmentation, and a consistent patching and vulnerability-management cadence
  • Operate our PCI-scoped environment to its required standard: change control, access reviews, audit evidence, and explicit lead approval for production changes - maintained throughout the migration
  • Enforce the data-residency requirement that PII stays in the EU region, through both policy and technical controls, with compliance treated as an integral part of migration planning
  • Ensure infrastructure leaving service is decommissioned securely: credentials rotated, access revoked, and data destruction documented
Requirements
  • Proven experience leading an SRE, platform, or infrastructure team, including running on-call rotations, acting as incident commander, managing performance, and developing engineers into senior roles
  • Deep production Kubernetes experience including self-managed or bare-metal clusters: cluster lifecycle, upgrades, networking (CNI, ingress, load balancing), and multi-tenant isolation
  • Strong GCP experience - GKE, IAM, VPC networking, Cloud SQL - ideally including a region or cross-region migration with production traffic
  • Senior-level Infrastructure as Code and GitOps experience - Terraform, Helm, and pipeline-as-code CI/CD; experience migrating away from legacy configuration management (Puppet, Ansible) is directly relevant
  • Strong observability and reliability practices - Prometheus, Grafana, SLOs, error budgets, and a track record of documentation and runbooks that prevent repeat incidents
  • Security-minded operations - least-privilege access design, secrets management (Vault or equivalent), network policy, and patching discipline, with security treated as a core part of reliability
  • Experience with a major infrastructure transition - a data centre migration, cloud migration, or platform rebuild with production traffic
  • 8+ years in production infrastructure, at least 2 in a lead or management role, with accountability for business-critical systems
  • Clear written and verbal English; based in the Greater Toronto Area and comfortable working across Toronto, Pune, Kraków, and Stockholm time zones
Additional Experience That Would Be Considered An Asset
  • OpenStack operations experience - Neutron networking, Cinder/Ceph storage, Octavia load balancing, Keystone identity
  • Talos Linux or another immutable, API-managed Kubernetes OS in production
  • Concourse CI or comparable pipelines-as-code platforms at organizational scale
  • Experience operating PCI-scoped or similarly regulated environments, including change control and audit discipline
  • Experience running infrastructure under a shared responsibility model with a parent company, partner, or major vendor, including cross-organization coordination
  • Regular use of agentic coding tools (Claude Code, Gemini, or equivalent) in your workflow, with sound judgement about validating AI-generated configuration before production use

Nice to have
  • Operational exposure to data systems our SRE team supports - Druid, Redpanda, Elasticsearch, Airflow
  • Network engineering depth - interconnects, BGP, site-to-site VPN, cross-region peering
  • GDPR data-residency, SOC 2, or ISO 27001 experience - the EU migration makes this increasingly relevant
  • Exposure to travel, flights, or large-scale search and cache systems
Compensation:
Canada - Toronto Office : 140, 000 - 180,000 CAD / Annual

Our pay ranges reflect the minimum and maximum target for new hire pay for the full-time position determined by role, level, and location.The pay range shown is based on our compensation structure in place at the time of posting and may be updated periodically based on business needs. Individual pay is based on additional factors including job-related skills, experience, and relevant education and/or training.

The targeted pay range listed reflects the base pay only and does not include bonus, or other benefits.

We use AI in our hiring process.

Similar Jobs

More Jobs at Etraveli Group

More Information Technology Jobs

Find similar Lead, Site Reliability Engineer jobs: