Barracuda Networks

Manager, Cloud Services and Site Reliability

Barracuda Networks$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years experience in SRE, DevOps, or technical operations including team leadership.
  • Strong grasp of cloud platforms, distributed systems, and reliability practices.
  • Experience in enhancing incident response, SLOs, SLIs, and monitoring.
  • Proven track record in hiring, mentoring, and developing engineers.
  • Excellent communication skills tailored to various stakeholders.
  • Ability to leverage data and structured problem solving for service reliability enhancement.
  • Knowledge of automation, CI/CD, disaster recovery, and multi-cloud operations.

Responsibilities

  • Lead and develop a high-performing Site Reliability Engineering team.
  • Drive implementation of reliability engineering practices across services.
  • Partner with engineering teams to enhance cloud system design and scalability.
  • Oversee incident management and continuous improvement post-incident.
  • Promote automation and tools that elevate operational consistency.
  • Utilize operational data to identify gaps and prioritize service improvements.
  • Ensure secure and compliant operations by collaborating with security teams.

Benefits

  • Equity through non-qualifying options.
  • Comprehensive health benefits.
  • Retirement plan with employer matching.
  • Opportunities for career growth and cross-training.
  • Flexible paid time off and vacation benefits.
  • Volunteer opportunities to give back to the community.
Full Job Description
Job ID: 27-0399

This is a hybrid position based in Ottawa, Ontario.

We are looking for a Manager, Site Reliability Engineering to lead a team responsible for the reliability, availability, scalability, and operational excellence of high-volume, business-critical SaaS applications. This role combines people leadership with strong technical judgment, helping the team improve reliability practices, reduce operational toil, and support resilient customer-facing services.

In this role, you will manage and develop SRE talent, partner closely with engineering, product, platform, and security teams, and help drive measurable improvements in service health, incident response, automation, and operational readiness. The application portfolio includes products such as Email Security Gateway and Cloud Email Archiving.
What You'll be Working on:
  • Lead, coach, and develop a high-performing SRE team, setting clear expectations, supporting career growth, and fostering a culture of ownership, collaboration, and continuous improvement.
  • Drive reliability engineering practices across critical services, including SLOs, SLIs, monitoring, alerting, capacity planning, and service health reporting.
  • Partner with engineering and platform teams to improve the design, operation, and scalability of cloud-based systems, with a focus on reliability, resilience, and maintainability.
  • Own and improve incident management practices, including major incident coordination, post-incident reviews, follow-up actions, and systemic reliability improvements.
  • Champion automation and tooling that reduce manual effort, improve operational consistency, and help the team scale support for production services.
  • Use operational data, service metrics, and risk indicators to identify reliability gaps, prioritize improvements, and communicate progress to technical and business stakeholders.
  • Support secure and compliant operations by partnering with security and engineering teams to embed appropriate controls, documentation, and operational practices into service delivery.
What you Bring to the Role:
  • 5+ years of experience in SRE, DevOps, infrastructure, cloud operations, or a related technical operations discipline, including experience leading or managing technical teams.
  • Strong understanding of cloud platforms, distributed systems, production operations, and modern reliability practices.
  • Experience implementing or improving SLOs, SLIs, monitoring, alerting, incident response, and post-incident review practices.
  • Demonstrated ability to hire, mentor, coach, and develop engineers while building a healthy, accountable, and inclusive team culture.
  • Strong communication skills, with the ability to explain technical topics clearly to engineering partners, product stakeholders, and business leaders.
  • Track record of using data, operational insight, and structured problem solving to improve service reliability and team effectiveness.
  • Experience with infrastructure automation, CI/CD practices, disaster recovery, cost optimisation, or multi-cloud operations.
  • Experience influencing operational change across teams, improving documentation practices, or evaluating tools and vendors that support service reliability.
What you'll get from us:

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility - there are opportunities for cross training and the ability to attain your next career step within Barracuda.
  • Equity, in the form of non-qualifying options
  • High-quality health benefits
  • Retirement Plan with employer match
  • Career-growth opportunities
  • Flexible Time Off and Paid Time Off benefits
  • Volunteer opportunities

At Barracuda, we believe in fair and equitable compensation practices that reflect both market realities and the unique circumstances of each geographical location. We recognize that cost-of-living disparities, market conditions, and other factors can significantly impact compensation expectations in different regions. The compensation range provided in this job description is for illustrative purposes only and may not reflect the actual compensation offers for the position in your location. Final compensation will be determined based on a variety of factors including the candidates' qualifications and experience.

#LI-Remote

About Barracuda Networks

Barracuda Networks is a provider of cloud-enabled security and data protection solutions for businesses. The company was founded in 2003 and is headquartered in Campbell, California. Barracuda Networks offers a range of products, including firewalls, email security, network security, and data protection solutions. The company's solutions are designed to protect against cyber threats, including malware, ransomware, and phishing attacks. Barracuda Networks serves customers in a variety of industries, including healthcare, finance, and education.
Learn more about Barracuda Networks
Size
1,500 employees
Industry
Founded
2003

Similar Jobs

More Jobs at Barracuda Networks

More Information Technology Jobs

Find similar Manager, Cloud Services and Site Reliability jobs: