About the Role We are hiring a Staff Site Reliability Engineer to raise the bar on how reliably KEV's platforms run in production, from the SLOs and error budgets that define what "reliable" means for each service to the incident practice that keeps issues from repeating. This role requires deep hands-on SRE experience and the judgment to assess where reliability is at risk, shape a strategy for closing the gap, and break that strategy into work a team can execute.
You will partner closely with KEV's DevOps Engineering team and product engineering teams to embed reliability into how services are built and operated, and coach and mentor other engineers on reliability practices. This is a role for someone who can surface trade-offs clearly and communicate them well.
How You Will Contribute- Reliability Strategy & SLOs Assess the current state of reliability across KEV's platforms, define and evolve the SLOs, SLIs, and error budgets that make "reliable" a measurable target rather than a feeling, and build a strategy for closing the gaps you find. Break that strategy into concrete workstreams and lead a team through execution.
- Incident Management & Postmortem Practice Own how KEV responds to production incidents, driving blameless postmortems that get to root cause rather than surface symptoms. Use recurring incidents and operational issues to identify durable fixes and longer-term reliability investments, not just one-off patches.
- Observability & Monitoring Platform Design and evolve the monitoring, alerting, logging, and tracing platform that gives engineering teams real visibility into system health. Use data to reduce noisy or ineffective alerts and make sure the signals teams rely on are ones they can trust.
- Automation & Toil Reduction Identify repetitive, manual operational work and replace it with automation and tooling that's maintainable, testable, and treated as production code. Build the tools yourself where that's the right answer, rather than only advocating for automation from the sidelines.
- Capacity & Performance Engineering Lead capacity planning and performance analysis so that KEV's platforms scale ahead of demand rather than in reaction to it. Use load testing and performance data to anticipate where systems will break before they do.
- AI-Assisted Reliability Operations Evaluate and apply AI tooling where it can meaningfully improve reliability and operational workflows - such as anomaly detection or incident triage - applying the same trade-off thinking to new tooling that you'd apply to any other operational decision.
- Coaching & Reliability Mentorship Coach and mentor other engineers on reliability practices - on-call hygiene, incident response, SLO-driven prioritization - raising the operational maturity of the broader engineering organization over time.
- Cross-Functional Partnership Partner closely with KEV's DevOps Engineering team and product engineering teams to embed reliability into how services are designed, built, and operated. Communicate trade-offs clearly so stakeholders understand what's changing, why, and what was weighed to get there.
Who You Are- 10+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, or a similar infrastructure/operations role, with a track record of leading reliability strategy and incident practice at scale.
- Familiarity working with Microsoft Azure and .NET, including legacy .NET Framework applications and IIS, so reliability practices hold up across KEV's full stack and not just its newest services.
- Demonstrated experience defining and operating SLOs, SLIs, and error budgets, and using them to guide engineering and operational priorities.
- Experience owning incident management processes and driving blameless postmortems that lead to durable fixes rather than one-off patches.
- Hands-on experience designing and building observability platforms - monitoring, alerting, logging, and tracing - that give engineering teams real visibility into system health.
- Strong automation skills, with experience building tools and scripts that eliminate repetitive manual operational work and are treated as production code.
- Experience with capacity planning and performance engineering, anticipating and preventing reliability issues rather than only reacting to them.
- Experience evaluating and applying AI tooling to improve reliability and operational workflows.
- Demonstrated ability to coach and mentor engineers on reliability practices, raising the operational maturity of a team over time.
- Strong ability to surface trade-offs and communicate them clearly to both engineering and cross-functional stakeholders.
What We Offer - Competitive compensation - We believe in rewarding great work with fair, competitive pay.
- Meaningful benefits- Because your well-being matters; both at work and at home.
- Retirement Savings Support - We help you plan for your future with company-matched programs, including RRSP matching in Canada and 401(k) contributions in the U.S.
- Professional development - We invest in your growth with ongoing learning, stretch opportunities, and continuing education, including KEV Academy for onboarding and skill-building, plus KEV University, our online platform offering a wide range of courses.
- Hybrid model - 3 days in the office to collaborate and connect, with flexibility the rest of the week.
- Flexible PTO - Take the time you need to recharge with close to 4 weeks of vacation and a company-wide holiday closure
- Office perks - Enjoy a fully stocked snack bar and occasional catered lunches-because we know that great conversations (and ideas) often start around good food.
Salary Range - $150,000-180,000
Note: We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar stage growth companies. Final offer amounts are determined by multiple factors, including skills, depth of work experience and relevant licenses/credentials, and may vary from the amounts listed above.