The role:Tyk Cloud is our managed API management platform, running on multi-region Kubernetes at scale for customers around the world.
We're looking for an SRE who's as comfortable in the code as in the infrastructure. You'll spend most of your time improving and automating the platform, and when you're on call, you'll handle incidents independently. You'll join a small, centralised SRE team that works closely with our product teams.
You don't need to have done everything below. We care most about how you reason through problems and how quickly you learn. You'll have three to four months to get up to speed, shadowing first, before you go on call.
What you'll do:Most of your time- Deliver the team's planned work each quarter, such as optimising the platform, building self-serve tooling for other teams, and rearchitecting parts of the platform as it grows.
- Help expand Tyk Cloud across regions and clouds, and bring down what it costs to run.
- Automate operations in Go, including building and maintaining our custom Kubernetes operators.
- Run the platform's services and databases, including MongoDB and Redis.
- Improve our observability: find the metrics that matter, and build the dashboards and alerts to act on them.
- Keep runbooks and documentation current, and support security work such as SOC 2 audits.
When you're on callYou'll be doing one week in three initially (one in four as we grow), Monday-Friday, on a 12-hour shift with secondary backup support. Rotas: 14:00-02:00 UTC
- Be first line for platform alerts and incidents: restore service, escalate or help fix product bugs, and lead post-incident reviews.
- Act as second line for our Customer Success team, on requests that come directly from customers.
- Handle ad hoc requests from other teams across the organisation regarding Tyk Cloud.
RequirementsWhat you'll need:- 3+ years in SRE, platform or infrastructure roles, across more than one company or production platform.
- Experience owning on-call and leading incidents yourself.
- Hands-on experience running production Kubernetes at scale, ideally EKS: operating, upgrading and debugging large, multi-tenant clusters.
- Experience designing and operating infrastructure on AWS, with Terraform or similar.
- The ability to write, test and ship Go tooling or services.
- Experience with Prometheus and Grafana, and with logging systems.
- Solid Linux and networking fundamentals (DNS, TCP/IP, HTTP, TLS, load balancing).
- Clear communication across time zones and teams.
Our stackEKS, Terraform/Terragrunt, Helm, GitHub Actions, Argo CD, MongoDB, Redis, Prometheus and Grafana.
BenefitsHere's why you should join us:- Everyone has unlimited paid holiday.
- We have total flexibility in hours, as we believe creativity flows better when our people are given freedom to decide when they are most productive. Everyone is unique after all.
- Employee share scheme
- Generous maternity and paternity leave
- Company retreats
What's it like to work here?! check it out: https://tyk.io/worklife/
You can see more about us here https://tyk.io