Site Reliability EngineerDepartment: Infrastructure
Employment Type: Full Time
Location: Austin
Reporting To: SRE Manager
DescriptionAs a Site Reliability Engineer at Brivo, you will bridge the gap between development and operations. You will ensure our global platform remains highly available while designing resilient systems and creating automation to eliminate operational toil. This is a hands-on role for someone who thrives in a collaborative, fast-paced environment and is eager to take ownership of platform reliability at a massive scale.
Responsibilities- Infrastructure Management: Build and maintain reliable, automated infrastructure across our cloud environments using Infrastructure as Code (IaC) tools.
- Incident Response: Participate in the on-call rotation, assisting with troubleshooting, root-cause analysis, and follow-up actions to prevent recurring incidents.
- Automation: Utilize strong scripting skills (Python, Bash, or Golang) to drive automation and reduce manual "toil".
- Observability: Apply best practices for monitoring and alerting using tools such as Prometheus/VictoriaMetrics and Grafana.
- Collaboration: Work with cross-functional partners to define meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs).
- Deployment Support: Support production readiness and contribute to the improvement of CI/CD tooling to empower our application teams.
Qualifications- Experience: 5+ years of experience as an SRE or in a related infrastructure-focused role.
- Systems: Strong experience managing Linux systems in production environments.
- Orchestration: Good working knowledge of Kubernetes or other container orchestration systems.
- Scripting: Solid abilities in Python or Bash; familiarity with Golang is a plus.
- Problem Solving: Proven ability to identify reliability issues and implement scalable improvements.
- Response: Hands-on experience with incident response and basic SLO/error-budget usage.
- AI Tooling: Familiarity with using or building internal tools that leverage LLMs to automate repetitive infrastructure tasks or incident response.
Individual compensation packages are based on job-related skills, experience, qualifications, work location, training, and market conditions. In addition, Brivonians enjoy a robust benefits and perks package tailored to their work location.