Barracuda Networks

Cloud Site Reliability Engineer II

Barracuda Networks • $100K — $120K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2-4 years of experience with AWS and/or Azure cloud infrastructure.
  • 1-2+ years hands-on with Kubernetes (EKS/AKS) containerized workloads.
  • Experience building dashboards in observability tools like Grafana.
  • Knowledge of Infrastructure as Code with Terraform and GitOps practices.
  • Scripting skills in Python or Bash; Go is a plus.
  • Curiosity to leverage AI coding tools in engineering workflows.
  • Strong analytical and communication skills, with a team-oriented mindset.

Responsibilities

  • Operate and scale LGTM telemetry infrastructure using Grafana, Loki, Mimir, and Tempo.
  • Design impactful Grafana dashboards for platform services and workloads.
  • Establish alerting strategies and SLO/SLI tracking to prevent production issues.
  • Automate deployment of monitoring agents across multi-cluster environments with GitOps.
  • Contribute to MTK Kubernetes platform health, performance tuning, and infrastructure updates.
  • Use AI coding tools to enhance automation and operational scripting.
  • Collaborate with internal teams on observability and performance troubleshooting.

Benefits

  • Opportunities for internal mobility and cross-training.
  • Equity in the form of non-qualifying options.
  • Comprehensive health benefits.
  • Retirement plan with an employer match.
  • Flexible Time Off and Paid Time Off.
  • Volunteer opportunities.
Full Job Description
As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda's next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.

Our team values collaborative knowledge sharing, automated reliability, and modern engineering practices. We actively embrace AI-assisted workflows (Claude Code, OpenCode, Codex CLI) to accelerate development, diagnostics, and routine platform maintenance. In this role, you will build and operate centralized observability stacks (Grafana, Loki, Mimir, Tempo), create actionable dashboards, automate telemetry via GitOps and Terragrunt, and partner with internal engineering teams to optimize application reliability in production.

Tech Stack Exposure:
  • Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager
  • Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize
  • IaC & Cloud Provisioning: Terragrunt, Terraform
  • GitOps & CI/CD: ArgoCD, GitHub Actions
  • Public Clouds: Amazon Web Services (AWS), Microsoft Azure
  • Languages & Scripting: Python, Bash (Go is a plus)
  • AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot

What you'll be working on
  • Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).
  • Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
  • Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.
  • Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.
  • Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.
  • Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.
  • Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.

What you bring to the role
  • 2-4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.
  • 1-2+ years of hands-on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
  • Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).
  • Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).
  • Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).
  • Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.
  • Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.

What you'll get from us

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility - there are opportunities for cross training and the ability to attain your next career step within Barracuda.
  • Equity, in the form of non-qualifying options
  • High-quality health benefits
  • Retirement Plan with employer match
  • Career-growth opportunities
  • Flexible Time Off and Paid Time Off benefits
  • Volunteer opportunities

What you'll get from us:

A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility - there are opportunities for cross training and the ability to attain your next career step within Barracuda. In addition, you will receive equity, in the form of non-qualifying options.

The anticipated salary range for this role is $100,000 CAD to $120,000 CAD. Actual compensation offered will be dependent upon the individual's skills, experience, and qualifications as they directly relate to the requirements of the position, the budget for the position, and applicable employment law.

Location: Ottawa, ON

#LI-hybrid

27-0459(a)

About Barracuda Networks

Barracuda Networks is a provider of cloud-enabled security and data protection solutions for businesses. The company was founded in 2003 and is headquartered in Campbell, California. Barracuda Networks offers a range of products, including firewalls, email security, network security, and data protection solutions. The company's solutions are designed to protect against cyber threats, including malware, ransomware, and phishing attacks. Barracuda Networks serves customers in a variety of industries, including healthcare, finance, and education.
Learn more about Barracuda Networks
Size
1,500 employees
Industry
Founded
2003

Similar Jobs

More Jobs at Barracuda Networks

More Information Technology Jobs

Find similar Cloud Site Reliability Engineer II jobs: