ServiceNow

Staff Software Engineer - SRE & AIOps

ServiceNow$125K — $220K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in software engineering or infrastructure operations, with significant SRE or DevOps experience.
  • Kubernetes expertise with proven history of managing large-scale production clusters.
  • Strong foundation in Infrastructure-as-Code practices with tools like Terraform or CloudFormation.
  • In-depth experience with cloud platforms such as AWS, Azure, and GCP relevant to SRE operations.
  • Proficiency in Linux systems and scripting languages such as Python, Go, or Bash.

Responsibilities

  • Design and operate enterprise-scale Kubernetes clusters across multi-cloud environments.
  • Implement automated systems for incident remediation to minimize downtime and manual intervention.
  • Develop and enhance SRE tooling and observability integrations to support global on-call operations.
  • Establish frameworks for service level objectives and alert policies, optimizing incident management.
  • Architect hybrid cloud operations and strategies for workload migration and disaster recovery.

Benefits

  • Flexible spending accounts and health plans.
  • 401(k) Plan with company match.
  • Employee Stock Purchase Plan (ESPP).
  • Flexible time away plan and family leave programs.
Full Job Description
Job Description

About the Role

ServiceNow is seeking a Staff Software Engineer - SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.

This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil across follow-the-sun global teams.

What you get to do in this role:
  • Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices that support high-velocity application deployments at 99.99%+ availability targets.
  • Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden.
  • Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data-driven incident response.
  • Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously.
  • Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails.
  • Architect hybrid cloud and data center operations, spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi-region deployments.
  • Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
  • Design on-call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post-incident review processes that capture learning and drive systemic improvements.
  • Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
  • Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams, establishing knowledge-sharing practices and technical documentation standards.
  • Reduce operational toil through systematic automation of repetitive tasks, from infrastructure provisioning to incident response to cost optimization, directly improving team capacity and job satisfaction across globally distributed operations.


Qualifications

To Be Successful in This Role You Have
  • Kubernetes Mastery: Strong hands-on expertise operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues.
  • Incident Auto-Remediation Expertise: Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms, that measurably reduce MTTR and on-call burden.
  • Cloud Platform Experience: Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region solutions.
  • DevOps & IaC Proficiency: Strong experience with Infrastructure-as-Code tools and GitOps platforms to drive reproducible, auditable infrastructure deployments.
  • SRE Tooling Fluency: Strong working knowledge of observability platforms, incident management systems, and log aggregation.
  • Distributed Systems Thinking: Solid understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, and proven ability to design systems resilient to these conditions.
  • On-Call Operations: Experience operating in follow-the-sun, 24/7 on-call models; ability to design escalation policies, runbooks, and communication patterns that balance responsiveness with operator well-being.
  • Data Center & Hybrid Cloud Operations: Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies.
  • AI/ML Integration: Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation.
  • Technical Leadership: Proven ability to drive technical decisions across teams and mentor engineers on reliability practices through credibility and technical depth.

Qualifications
  • Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
  • 8+ years in software engineering or infrastructure operations, with 5+ years in SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience.
  • 4+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at scale.
  • Proficiency in Infrastructure-as-Code: Terraform, CloudFormation, or equivalent tools used to manage infrastructure at scale.
  • Public Cloud Expertise: Demonstrable experience across 2+ of the following: AWS, Azure, GCP, with solid knowledge of services relevant to SRE operations (compute, networking, storage, observability).
  • On-Call Operations: Experience operating or designing components of 24/7 follow-the-sun on-call models for distributed teams, including runbook development and incident response.
  • Incident Auto-Remediation: Proven ability to design and implement automated remediation systems that measurably reduce manual toil.
  • Linux & Systems Programming: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
  • SRE Mindset: Demonstrated commitment to reliability through engineering, favoring durable automation over heroics, and data-driven decision-making.
  • Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience).

Preferred:
  • Kubernetes certification (CKA, CKAD, or equivalent).
  • Experience with service mesh platforms or advanced networking in Kubernetes environments.
  • Background in migrating workloads from on-premises data centers to public cloud environments.
  • Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing).
  • Track record of mentoring infrastructure engineering teams.

Why This Role?

This role offers the opportunity to eliminate operational toil at scale and build the automation-first infrastructure practices that define modern cloud operations. You will architect systems that allow ServiceNow's global teams to operate reliably with confidence in automated remediation systems actively preventing and resolving failures. Your work will directly shape how the organization scales reliability as it grows, establishing patterns that benefit teams across the platform. This is a role for an engineer who wants technical depth, team influence, and the satisfaction of watching systems recover from failure automatically.

For positions in this location, we offer a base pay of C$125,700 - $220,000, plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as qualifications, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.

Additional Information

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

About ServiceNow

ServiceNow provides cloud-based solutions that define, structure, manage, and automate services for enterprise operations in North America, Europe, the Middle East, Africa, the Asia Pacific, and other countries. The company offers service management solutions, including incident, problem, change, request, and cost management as well as service catalogs; and IT, HR, facilities, and field service management solutions. It also provides IT operations management solutions covering service mapping, delivery, and assurance solutions; business management solutions such as financial management, project portfolio suite, vendor performance management, and performance analytics as well as governance, risk, and compliance; and application development services.

ServiceNow Careers

Join the dynamic team at ServiceNow, a global leader in digital workflow solutions, where innovation and leadership converge to shape the future of work. At ServiceNow, we offer more than just job opportunities; we provide a platform for professional growth and a chance to be part of a culture that values diversity, creativity, and continuous learning.

Work You’ll Do

Embark on a career journey with ServiceNow and contribute to the world’s leading enterprises' digital transformation. Our team is at the forefront of developing cutting-edge technologies that improve how people work. With ServiceNow, you will use your skills to impact businesses and industries profoundly, driving efficiency and innovation.

Join Our Market-Leading Team

ServiceNow is not just another technology company. We are a team that thrives on diversity and leadership, fostering an inclusive environment that promotes growth and development. Our commitment to diversity training ensures that every team member can achieve their potential.

Innovative Work

ServiceNow is home to more than 10,000 dedicated professionals who lead the charge in digital workflows and enterprise solutions. As part of our team, you will engage in projects that merge technology with practical applications, creating revolutionary products that advance how services are delivered and managed.

Career Development

At ServiceNow, your career trajectory is filled with boundless opportunities. We support your growth with robust training programs, leadership development courses, and access to global challenges. Whether you are looking for an internship, full-time position, or leadership role, ServiceNow equips you with the tools to excel.

Be Part of a Great Team

Working at ServiceNow means being part of a community that values teamwork and innovation. Our collaborative environment encourages networking and sharing ideas, making our workplace vibrant and dynamic. The benefits of joining ServiceNow extend beyond comprehensive health and wellness; they include fostering professional connections and friendships that last a lifetime.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional or a recent graduate, ServiceNow offers a range of employment options to suit your career goals. From internships that provide real-world experience to full-time positions that challenge you to leverage your expertise, we are committed to hiring the best talent.

Stay Connected

Join Our Team Search open positions that match your skills and interests. At ServiceNow, we look for passionate, curious, and solution-driven team players. Explore the possibilities that await you at a company that is committed to your professional success.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at ServiceNow.

ServiceNow Careers

Empowering professionals to achieve more, ServiceNow is where careers are future-proofed, and ambitions are realized. Join us in our journey of growth and innovation.
Learn more about ServiceNow
Size
16,881 employees
Market Cap
$76.5 billion
Industry
Net Income
$118.5 million
Founded
2004
5 Year Trend
+33.5%
Revenue
$4.5 billion
NASDAQ

Similar Jobs

More Jobs at ServiceNow

More Information Technology Jobs

Find similar Staff Software Engineer - SRE & AIOps jobs: