Backblaze

Strategic Operations Engineer III

Backblaze$123K — $175K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in IT Operations, SRE, or similar roles.
  • Strong expertise in Incident, Problem, and Change Management (ITIL or similar frameworks).
  • Proven experience in governing and optimizing operational processes.
  • Strong knowledge of AI/ML concepts, including anomaly detection, predictive analytics, and data modeling.
  • Hands-on experience with AIOps platforms or building AI-driven operational solutions.

Responsibilities

  • Lead and govern the end-to-end incident management lifecycle.
  • Drive major incident management processes and communications.
  • Improve Mean Time to Resolution (MTTR) through automation and process optimization.
  • Maintain and improve intelligent heatmaps using AI/ML for recurring technical theme identification.
  • Implement trend analysis for proactive problem identification using observability data.
  • Govern change management processes, ensuring safe and compliant deployments.
  • Leverage AI/ML for anomaly detection and automated root cause analysis.

Benefits

  • Culture of learning and development.
  • Encouragement for diverse self-expression.
  • Collaborative team environment.
Full Job Description
What You'll Do:
Incident Management
  • Available to Lead and govern the end-to-end incident management lifecycle, including detection, triage, escalation, and resolution.
  • Drive major incident management (MIM) processes and communications.
  • Improve MTTR (Mean Time to Resolution) through automation and process optimization.
  • Establish and maintain incident response playbooks and runbooks.
Problem Management
  • Maintain and improve intelligent heatmaps leveraging AI/ML to identify recurring technical themes and prioritize long-term remediation.
  • Implement trend analysis and proactive problem identification using observability data and AI.
  • Track and manage problem records to closure.
Change Management
  • Govern change management processes (lead the CAB), ensuring safe, compliant, and low-risk deployments.
  • Define and enforce change policies, risk assessments, and approval workflows.
  • Drive continuous improvement in release and deployment practices.
Observability & Service Reliability
  • Maintain a strong understanding of system architecture and monitoring strategies, identifying gaps and opportunities for improvement.
  • Partner with engineering teams to improve system resilience and performance.
  • Reduce alert fatigue by improving signal-to-noise ratio in monitoring systems.
AI-Driven Operations (AIOps)
  • Leverage AI/ML for anomaly detection, predictive alerting, and automated root cause analysis.
  • Implement AI-driven solutions to optimize incident response and operational workflows.
  • Analyze large-scale operational data to identify patterns and recommend improvements.
  • Experience with AIOps platforms or building AI-driven operational solutions.

Required Qualifications:
  • 5+ years of experience in IT Operations, SRE, or similar roles.
  • Strong expertise in Incident, Problem, and Change Management (ITIL or similar frameworks).
  • Proven experience in governing and optimizing operational processes.
  • AI & Data Expertise: Strong knowledge of AI/ML concepts, including anomaly detection, predictive analytics, and data modeling.
  • AIOps Experience: Hands-on experience with AIOps platforms or building AI-driven operational solutions (event correlation, alert prioritization).

Preferred Qualifications:
  • ITIL certification (Foundation or higher).
  • Proficiency with platforms such as Jira, SNOW, FireHydrant, Moogsoft, etc.
  • Experience working in high-availability, large-scale environments.

Key Competencies:
  • Positive Attitude!
  • Strong analytical and problem-solving skills.
  • Process-oriented mindset with a focus on governance and continuous improvement.
  • Excellent stakeholder communication and leadership skills.
  • Ability to drive change across cross-functional teams.

At this point, we hope you're feeling excited about the job description you're reading. Even if you don't meet every requirement, we still encourage you to apply. Learning, developing, and growing are key parts of our culture. We're eager to meet people who believe in our mission and can contribute to our team in various ways. We want people to feel comfortable expressing their true selves and to come, stay, and do their best work here.

To provide greater transparency to candidates, we share base pay ranges for all US-based job postings regardless of state. We set standard base pay ranges for all roles based on function, level, and country location, benchmarked against similar-stage growth companies. Final offer amounts are determined by multiple factors, including candidate location, skills, depth of work experience, and relevant licenses/credentials, and may vary from the amounts listed below.

The expected salary range for this role is $123,000 - $175,000.

About Backblaze

Backblaze is a data storage company that provides cloud storage and backup solutions for businesses and consumers. The company was founded in 2007 and is headquartered in San Mateo, California. Backblaze offers a range of products, including cloud storage, backup software, and data migration services. The company's cloud storage service is designed to be affordable and easy to use, with no hidden fees or complicated pricing structures. Backblaze has over 1 million customers and stores over 1 exabyte of data.
Learn more about Backblaze
Size
200 employees
Market Cap
$165.8 million
Industry
Founded
2007
NASDAQ

Similar Jobs

More Jobs at Backblaze

More Information Technology Jobs

Find similar Strategic Operations Engineer III jobs: