Location / Timezone: Remote, between Central Europe and US East timezones for optimal collaboration with the team.
Description of the RoleAs a Senior Site Reliability Engineer, you will be an early contributor to our dedicated SRE function, reporting directly to Florian Beer. This is a high-impact, autonomous role where you will design and implement the systems that support Laravel Cloud, Nightwatch, and Forge. You will act as a bridge between development and operations, advocating for a blameless culture and shared responsibility for reliability across the entire organization.
Your 12-Month Mission Imagine we are all at Laracon in 12 months' time. You are telling the team about your first year, and the impact is undeniable:
- First 30 Days: You will have selected an SLO monitoring solution.
- Day 60: You will have guided teams towards identifying SLIs and how/why we work with SLOs at Laravel.
- Day 90: You have established clear, data-driven SLOs across engineering teams, giving us a unified language for reliability.
- Year One: You have worked together with our engineering teams to identify SLIs, review existing SLAs, created SLOs, and educated teams on how SLO's error budgets guide decisions.
What You Will Do- Architect Reliability: Establish SRE as a core function at Laravel, building the fundamentals from the ground up.
- System Design: Design, build, and maintain multi-region Kubernetes infrastructure and global distributed systems.
- Automation: Solve operational challenges through software, reducing manual intervention (toil) for our product teams.
- Observability: Design and implement monitoring, logging, and alerting systems using tools like Prometheus, Grafana, and Loki.
- Collaboration: Partner with product leads and SecOps to make reliability a shared responsibility.
RequirementsRequirements - What You Will Bring - Infrastructure Mastery: Deep experience with Linux system administration and cloud platforms, specifically AWS.
- Orchestration & IaC: Proficiency with Kubernetes, Docker, and managing infrastructure via Terraform.
- Programming Skills: The ability to solve problems with software and scripting using e.g. PHP, Bash, or Go.
- Systems Thinking: A "smart and passionate" approach to troubleshooting, with the ability to deconstruct complex systems into triagable components.
- Reliability Mindset: Experience with SLO/SLI/SLA definition, capacity planning, and performance tuning.
- Soft Skills: A commitment to documentation, cross-team collaboration, and an automation-first mindset.
Requirements - Bonus Skills- Framework Familiarity: Previous experience working with the Laravel framework and our existing product suite (Cloud, Forge, Vapor, etc.) is highly preferred.
- Advanced Observability: Experience with Prometheus, Grafana Mimir, and Grafana Loki for metrics storage and alerting.
Benefits- Small tight-knit team where every developer counts
- Fully remote and globally distributed working environment
- Option to attend Laracon conferences around the world
- Health care plan (Medical, Dental & Vision)
- Paid time off (Vacation, Sick & Public holidays)
- Family leave (Maternity, Paternity)
- Pension plans (As locally applicable)
- Performance based bonus plan
- Company equity