IBM

Staff Site Reliability Engineer - Confluent Incident Management & Reliability

IBM$120K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10+ years of experience in SRE, incident management, or reliability engineering
  • Cloud expertise with AWS, GCP, or Azure
  • Experience in reliability programs at large organizations (500+ engineers)
  • Deep knowledge of incident management tools (Rootly, PagerDuty)
  • Strong grasp of distributed systems and cascading failure modes
  • Kafka/event streaming expertise preferred
  • Bachelor's degree required; Master's preferred

Responsibilities

  • Analyze systemic failure patterns and design reliability improvements
  • Own configuration and workflows for disaster management tools like Rootly
  • Define and maintain SLO/SLA frameworks to guide reliability efforts
  • Establish standards for incident response and facilitate continuous improvement
  • Review customer-facing incident documents for quality and clarity
  • Develop and deliver training programs for incident management
  • Collaborate with engineering leaders to enhance reliability practices

Benefits

  • Global team with follow-the-sun coverage for work-life balance
  • Opportunity for hands-on engineering alongside strategic leadership
  • Engagement in evolving incident response practices
  • Role focused on proactive improvements to prevent incidents
  • Chance to influence reliability standards across a diverse engineering landscape
Full Job Description
Your role and responsibilities

About the Role:

Confluent Cloud processes millions of events per second across AWS, GCP, and Azure. When incidents happen in a multi-cloud streaming platform, they happen at scale-data in motion, exactly-once semantics, and cascading failure modes that require deep systems thinking. We need an expert-level engineer who can drive proactive reliability improvements that prevent these incidents before they occur.

This role combines hands-on technical work with strategic program ownership. You'll spend roughly 75% of your time on engineering: building automation, improving tooling, analyzing systemic failure patterns, and designing reliability improvements. The remaining 25% is teaching and coordination: coaching teams through post-mortems, training incident commanders, and evolving our incident response practices.

You'll be part of a global team with follow-the-sun coverage, with clean handoffs that keep everyone working sustainable hours. This role sits within Cloud Architecture and Reliability - Supportability, a horizontal team that owns reliability standards and tooling across engineering. You're the person who makes us need incident management less.

What You Will Do:
  • Analyze systemic failure patterns and design reliability improvements that prevent incident recurrence
  • Own Rootly configuration, workflows, and integrations with PagerDuty, Jira, Confluence, and Slack
  • Define and maintain SLO/SLA frameworks; use error budgets to guide reliability investments
  • Own standards, practices, and continuous improvement of incident response across engineering
  • Edit and review customer-facing incident documents (CRCAs) to ensure quality and clarity
  • Develop and deliver training programs; coach teams through post-mortems
  • Partner with engineering leaders to elevate reliability practices org-wide
  • Deep experience with observability: metrics, logging, tracing
  • Kubernetes and container orchestration experience
  • Understanding of CI/CD pipelines and release processes
  • Strong written communication (design docs, runbooks, post-mortems)
  • Experience driving org-wide process and cultural changes


Required education

Bachelor's Degree

Preferred education

Master's Degree

Required technical and professional expertise

  • 10+ years of relevant experience in SRE, incident management, or reliability engineering
  • Cloud experience with at least one of AWS, GCP, or Azure (we run all three)
  • Experience navigating reliability/incident programs at 500+ engineer organizations
  • Deep expertise with incident management tooling (Rootly, PagerDuty, or similar)
  • Strong understanding of distributed systems and failure modes at scale
  • Kafka/event streaming expertise preferred, or demonstrated rapid mastery of complex systems


Preferred technical and professional experience
• Advanced Cloud Knowledge: Experience with cloud-based infrastructure and its application in reliability and resiliency engineering. • Specialized Scripting Skills: Proficiency in scripting languages and automation tools to optimize system reliability and performance.

OTHER RELEVANT JOB DETAILS

Must have the ability to work in Canada without sponsorship.

This role will involve working with technology that is covered by Export Regulations sanctions. If you are a Foreign National from any of the following US sanctioned countries (Cuba, Iran, North Korea, Syria, and the Crimea, Luhansk, Donetsk, Kherson, and Zaporizhia regions of Ukraine) on a work permit, you are not eligible for employment in this position.

The salary range for the position is based on a full-time schedule. Your ultimate salary within this range may vary depending on your job-related skills and experience for this position.

About IBM

Sequent Computer Systems designs and manufactures multiprocessing computer systems based on Cache Coherent Non-Uniform Memory Access architecture, which are used primarily as application and database servers for commercial applications. Through a partnership with Oracle Corporation, Sequent became a dominant high-end UNIX platform in the late 1980s and early 1990s. Later it introduced a next-generation high-end platform for UNIX and Windows NT based on non-uniform memory access architecture, NUMA-Q. Sequent Computer Systems, Inc. was incorporated in 1983 and is based in Beaverton, Oregon. In July 1999 Sequent Computer Systems, Inc. was acquired by International Business Machines Corp.

IBM Careers

Joining IBM presents a prime opportunity to be part of our global team of professionals, leading the way in technological innovation and business solutions. At IBM, we offer more than just job opportunities; we provide a platform for growth, leadership development, and a chance to be at the forefront of industry innovation. Work You’ll Do At IBM, your work impacts global markets and helps reshape industries. As part of our team, you will contribute to projects that harness the power of cloud computing, AI, and blockchain technologies. With IBM, you are positioned to lead in the marketplace, leveraging our deep industry expertise and commitment to digital innovation. Transform Your Career IBM is not just a company; it's a culture of innovation and leadership. We are committed to diversity and providing an inclusive environment where all professionals can thrive. Joining IBM means being part of a team that values your unique skills and perspectives. IBM offers a variety of career paths, including full-time positions and internships, that allow you to explore your interests and develop new skills. Our professional development programs are designed to help you grow at every stage of your career, featuring robust training, certification support, and opportunities for advancement. Innovate with Us Engage in work that matters with a team of over 350,000 IBMers worldwide, driving progress in over 170 countries. At IBM, innovation isn’t just about technology, it’s about transforming how businesses operate and compete. Your work will deliver solutions that anticipate and fulfill the needs of our clients in an ever-evolving landscape. Be Part of a Great Team IBMers are diverse, talented, and globally connected. You’ll collaborate with thought leaders and industry experts who are shaping the future of technology. Our culture fosters creativity, teamwork, and the continuous pursuit of excellence. Networking and Professional Growth IBM is deeply invested in the professional growth of its employees. We encourage networking within the company to foster connections that can lead to greater innovation and career advancement. Our leadership is committed to providing every employee with the tools they need to succeed, from mentorship programs to diversity training. Explore Job Opportunities Whether you’re looking for an internship, a graduate role, or a leadership position, IBM offers a range of career opportunities. Our hiring process is designed to be transparent and fair, providing you with all the resources you need to succeed in your interview and resume preparation. Stay Connected Join Our Team Discover the various positions available that match your skills and interests. At IBM, we look for passionate, curious, and solution-driven team players. Explore our open positions and find where you can make a difference. Keep Up to Date Stay informed with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here. Job Alert Emails Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. See what exciting and rewarding opportunities await at IBM. Join IBM and be part of a legacy of leadership and innovation. Shape your future and build a career designed to last.
Learn more about IBM
Size
307,600 employees
Market Cap
$128 billion
Industry
Net Income
$5.5 billion
Founded
1911
5 Year Trend
-6.4%
Revenue
$73.6 billion
NASDAQ

Similar Jobs

More Jobs at IBM

More Information Technology Jobs

Find similar Staff Site Reliability Engineer - Confluent Incident Management & Reliability jobs: