Position descriptionAs a member of one of our engineering teams, you'll be a key player in making SonarQube Cloud and SonarQube Server the best tools for Code Quality and Security, providing new features to deliver high-quality and powerful products and services that help our users write better software.
You will have the opportunity to see your features come to life in Production with short iteration loops. While keeping our roadmap and business priorities in mind, you will be able to have a high impact on the software that we own and develop.
By joining us you will bring your experience and expertise to help push our product to its next stage of evolution and fulfill the needs of our large user and customer base.
What you will do- Lead and manage a team of Engineers (Cloud Engineers and Site Reliability Engineers), providing guidance, support, and mentorship to help individuals grow in autonomy and master the complexities of cloud operations.
- Hold yourself and your team accountable to high engineering standards, specifically focusing on system reliability, performance, and security.
- Manage the team's operational workload, including on-call health, incident response, and the reduction of manual toil through automation.
- Foster a safe culture of feedback and continuous improvement, encouraging the team to conduct blameless post-mortems and share architectural insights.
- Collaborate with other value stream squads to ensure the production platform meets the needs of our developers while maintaining strict governance.
- Communicate a clear vision for the squad that aligns with the platform engineering roadmap, focusing on resiliency, business continuity, and cost optimization.
- Partner with the Hiring team to recruit talented Engineers for the team. Participate in improving the hiring process for the team and ensuring we recruit enough to reach our goals.
- Lead by example, modeling behaviors of servant leadership and high-stakes decision-making.
Experience and qualifications- 10+ years of experience in software engineering, with a significant focus on Site Reliability Engineering (SRE), Cloud Operations, or Infrastructure Engineering.
- Operational Mindset: Deep understanding of the DevOps/SRE mindset, including experience managing mission-critical shared services (e.g., Aurora DBs, OpenSearch, Control Planes).
- Cloud Expertise: Advanced knowledge of AWS (or similar providers) and experience managing organizational-scale infrastructure (IAM, OUs, Account Vending).
- Infrastructure as Code: Experience with the coding lifecycle in an infrastructure context (e.g., Python, CDK, Terraform) and the ability to perform rigorous code reviews for infrastructure components.
- Observability & Resiliency: Proven experience defining observability patterns (logging, tracing, metrics) and designing Disaster Recovery/Business Continuity strategies.
- Agile & FinOps: Experience with Agile methodologies and a strong understanding of cloud cost optimization (Rightsizing, Spot instances, Reserved Instances).
Additional commentsAdditional commentsThis role is based in Austin TX. We are unable to consider candidates unwilling to be in Austin, but we are willing to relocate the right candidate.
Benefits- Flexible comprehensive employee benefit package.
- We encourage usage of our robust time-off allocations. You will receive 23 days of PTO per calendar year (on a pro-rated basis depending on your employment start date), with additional time provided for sickness, life events and holidays
- .We offer an exciting 401(k) plan that has a 4% match, fully vested on day one of participation.
- Fully paid parking in the heart of downtown Austin, Texas.
- Global workforce with employees in 20+ countries representing 35+ unique nationalities.
- We have an annual kick-off somewhere in the world where we meet to build relationships and goals for the company.
- Monthly catered events, and team events