Full Job Description
This role exists to make the correct way to build and ship at Klue the easy way. When the paved road, shared infrastructure, CI/CD, and guardrails are solid, every engineer at Klue gets faster, safer, and more cost effective to support, without having to think about it.
You'll work alongside the rest of the DevOps team, as well as engineering managers across Product, Data, and AI, and our Security and Compliance stakeholders.
This is a role for someone who is systems-minded, pragmatic, security-aware, and allergic to toil. Someone who'll jump in when the platform's on fire, but always ask what change stops the next one from starting. You balance proactive and reactive work, and when you get to choose, you choose proactive. That's what moves the ratio in the right direction.
What's In It for You?
You'll own real infrastructure at real scale and have deep, hands-on experience across cloud, platform engineering, infrastructure, and developer experience, with authority over the IaC, GKE, and CI/CD architecture every team depends on. You'll help shape how a fast-moving, AI-first engineering org ships software safely.
Klue moves fast and builds AI-first. This role involves keeping engineering fast, the platform scalable and secure, and costs fully visible. The four areas below are where most of that lives, not where it ends:
What You'll Do
- Own the paved road for infrastructure. Build and extend our Pulumi component library so teams get correct, disaster-recovery-ready infrastructure by default instead of hand-rolling stacks across our GCP environments.
- Run GKE like a product. Handle cluster and node pool upgrades, in-cluster services, workload right-sizing, and policy enforcement so Kubernetes stays reliable as we scale.
- Make CI/CD something engineers never think about. Own the core building blocks of our CI/CD workflows, cut CI time and flake rate fleet-wide, and standardize the golden path for shipping services.
- Give engineers real visibility into their own systems. Improve the internal developer platform surfaces that show what's actually happening under the hood.
- Keep production legible. Own the health of our observability stack: consistent instrumentation, actionable alerts, and less noise on-call.
- Treat cost like an engineering problem. Drive GCP and LLM cost reduction through attribution, right-sizing, commitment planning, and killing off orphaned resources.
- Make the secure path the fast path. Extend our software supply chain controls, own infrastructure identity and access, and build SOC 2 and other compliance requirements into automated, evidenced controls, so audits are a byproduct of how we build.
What Success Looks Like
We want you to love the role you are in and believe these skills and experiences will help you to be successful!
- Self-Serve Platform: New services are provisioned from standardized components without DevOps writing bespoke infrastructure
- Reactive Load Trends Down: DevOps spends less time on reactive work every quarter, on-call page volume moves in the same direction.
- CI Speed & Trust: Median pipeline duration and flaky-failure rate both drop measurably across the fleet
- Security & Audit Readiness: Controls are automated and evidenced, with audit findings in your areas trending to zero.
- Cost Efficiency: GCP spend per unit of workload declines through right-sizing and cleanup, with clear per-team attribution that makes tradeoffs visible to their owners.
- Reliability: The services that matter have published SLOs, real error budgets, and incident reviews that produce durable platform fixes.
- Automation Over Repetition: Recurring manual work gets automated, not repeated.
What You Bring
Must-Haves
- Deep, hands-on experience in DevOps, SRE, platform, or infrastructure engineering, operating production systems that paying customers depend on, including being on-call for them.
- Deep, practical Kubernetes knowledge.
- Infrastructure as code treated as software. Pulumi or Terraform.
- Strong GCP or AWS fundamentals.
- You have a background in systems engineering and writing software.
- CI/CD owned as a product. You've built and maintained pipelines used by other teams, GitHub Actions or Buildkite preferred, GitLab CI / CircleCI transfers.
- Hands-on experience with security tooling as an engineering practice.
- Compliance engineering for SOC 2 or enterprise partner security programs implemented as automated controls; threat modeling, etc
- Ownership under ambiguity. You scope a vague problem, ship it incrementally, and tell people clearly what changed and why.
- Fluent with AI coding tools, with experience working with AI-first teams.
Nice-to-Haves
- You have owned platform migrations and decommissions.
- Built an internal developer platform or self-service tooling that other engineers adopted.
- Comfort with LLM infrastructure (gateways, token cost attribution, rate limiting, provider failover).
Our Tech Stack
- GCP
- Kubernetes
- Pulumi IaC
- Postgres
- Temporal
Not ticking every box? That's okay. We take potential into consideration. An equivalent combination of education and experience may be accepted in lieu of the specifics listed above. If you know you have what it takes, even if that's different from what we've described, be sure to explain why in your application.
How We Work at Klue
We love the balance of connection and flexibility. We work Office-First, with ahybrid touch, and work together in our vibrant offices on Mondays, Wednesdays & Thursdays to collaborate, brainstorm, and build together.
Our main hiring hubs are in:
- C🇦 Vancouver (HQ)
- C🇦 Toronto
- GB London
Our Commitment to You
- High Performance Culture. We reward high performance and growth through career development, coaching, and annual performance reviews.
- Comprehensive benefits: Extended health & dental coverage that starts on Day 1. Fun perks like discounts at Goodlife and Perkopolis are gravy.
- Ownership: All full-time employees have the opportunity to participate in our Employee Stock Option Plan.
- Our Vacation Policy is Take The Time You Need. We just ask that you give notice and don't leave your team hanging.
- Top-tier tools. All employees will receive a Mac (or PC, if that's your jam) and access to A+ tooling.
- AI First. All employees are encouraged to lean into AI to work smarter and faster. Built something cool lately? Show us at our Friday Show, Don't Tell Meetings.
- Growth / Leadership. Direct access to our leadership team, including our CEO, and opportunities to connect with incredible people across the company.
- Social connection. There's no shortage of ways to stay connected and have fun. We get together once a year in Vancouver for a company- wide kickoff. Throughout the year our Hubs hold regular social events.
- Dog-friendly spaces. Bring your four-legged friend along in Vancouver or Toronto, as our offices are pup-approved.