Job SummaryWe're looking for a
Site Reliability Engineer (SRE) who's passionate about building resilient, high-performing systems that our customers can depend on every day. In this role, you'll blend software engineering and operations expertise to keep our Azure-based platforms fast, secure, and always online
You'll work side-by-side with developers, product teams, and DevOps to design scalable architecture, automate deployment processes, and drive performance improvements across our ecosystem. If you love problem-solving, reliability challenges, and working in a fast-paced environment where every improvement matters - this is the role for you
Job Responsibilities- Ensure uptime, reliability, and scalability across our cloud environments.
- Partner with development teams to design and run load, performance, and chaos testing.
- Build, monitor, and optimize systems within Microsoft Azure - including App Service Plans, Azure Functions, Event Grid, Event Hub, Service Bus, and Azure SQL Service.
- Manage incident response, ensuring issues are quickly resolved and documented for future prevention.
- Continuously improve observability, automation, and performance tuning to strengthen reliability.
- Lead post-incident reviews and drive long-term fixes for systemic issues
Job Requirements- 3-5 years of experience in a Site Reliability or similar engineering role.
- Deep knowledge of Azure cloud services.
- Proven experience monitoring, debugging, and supporting large-scale distributed systems.
- Strong background in CI/CD, performance diagnostics, and automation.
- A creative, analytical mindset and a drive to make complex systems perform better every day.
- Bachelor's degree in Computer Science, Engineering, or equivalent hands-on experience.
Bonus Points If You Have:- Experience with Azure DevOps, Agile teams, load testing, and chaos testing.
- A passion for continuous learning, improvement, and innovation.