ServiceNow

Manager, Network Reliability and Resiliency

ServiceNow$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in network engineering, reliability, or cloud infrastructure
  • Proven experience managing engineers with a focus on development and accountability
  • Technical expertise in troubleshooting across network services and Linux systems
  • Understanding of networking technologies (TCP/IP, routing, DNS) without needing deep expertise
  • Experience in managing significant incidents in high-pressure environments
  • Knowledge of SRE principles like SLIs, SLOs, and post-incident improvement
  • Background in automation and scripting to reduce operational workload
  • Familiarity with AI tools to enhance workflows and analysis

Responsibilities

  • Lead and coach a team of network reliability engineers
  • Oversee incident response and manage high-severity incidents
  • Engage in complex troubleshooting and customer escalations
  • Implement SRE practices to enhance network reliability and performance
  • Drive automation of repetitive and error-prone tasks
  • Collaborate with cross-functional teams to improve service reliability
  • Ensure operational readiness and adherence to operational acceptance criteria

Benefits

  • Opportunities for professional development and career growth
  • Access to innovative technologies and practices
  • Flexible work personas to promote work-life balance
  • Supportive team environment emphasizing coaching and development
  • Engagement with globally distributed teams for diverse perspectives
Full Job Description
Description du poste

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status. Employment is contingent upon successful completion and maintenance of the required screening.

What you get to do in this role:

We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.

You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.

Lead and develop the team
  • Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
  • Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
  • Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
  • Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.

Provide technical and incident leadership
  • Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
  • Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
  • Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
  • Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.

Improve reliability through SRE practices
  • Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
  • Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
  • Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
  • Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.

Embed automation in daily operations
  • Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
  • Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
  • Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
  • Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.

Partner across the organization
  • Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
  • Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
  • Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
  • Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.


Qualifications

To be successful in this role you have:
  • Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
  • Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability.
  • Sufficient hands-on technical background to guide production troubleshooting across Linux-based systems and network services. You can interpret logs, metrics, alerts, and packet-level evidence and make sound operational decisions.
  • Working knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not required.
  • Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment.
  • Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post-incident improvement.
  • An automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistency.
  • Experience working with geographically distributed teams and cross-functional partners in software, platform, infrastructure, or cloud services.
  • Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detail.
  • Experience using or evaluating AI-assisted tools to improve analysis, decision-making, automation, or team workflows, with appropriate attention to accuracy, security, and operational risk.

Preferred qualifications
  • Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment.
  • Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking.
  • Experience improving observability, change safety, capacity management, or operational readiness for production services.
  • Experience with IT service management practices, including incident, change, and problem management.
  • Relevant certifications such as CCNA, CCNP, Azure/AWS/GCP related


Informations complémentaires

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

About ServiceNow

ServiceNow provides cloud-based solutions that define, structure, manage, and automate services for enterprise operations in North America, Europe, the Middle East, Africa, the Asia Pacific, and other countries. The company offers service management solutions, including incident, problem, change, request, and cost management as well as service catalogs; and IT, HR, facilities, and field service management solutions. It also provides IT operations management solutions covering service mapping, delivery, and assurance solutions; business management solutions such as financial management, project portfolio suite, vendor performance management, and performance analytics as well as governance, risk, and compliance; and application development services.

ServiceNow Careers

Join the dynamic team at ServiceNow, a global leader in digital workflow solutions, where innovation and leadership converge to shape the future of work. At ServiceNow, we offer more than just job opportunities; we provide a platform for professional growth and a chance to be part of a culture that values diversity, creativity, and continuous learning.

Work You’ll Do

Embark on a career journey with ServiceNow and contribute to the world’s leading enterprises' digital transformation. Our team is at the forefront of developing cutting-edge technologies that improve how people work. With ServiceNow, you will use your skills to impact businesses and industries profoundly, driving efficiency and innovation.

Join Our Market-Leading Team

ServiceNow is not just another technology company. We are a team that thrives on diversity and leadership, fostering an inclusive environment that promotes growth and development. Our commitment to diversity training ensures that every team member can achieve their potential.

Innovative Work

ServiceNow is home to more than 10,000 dedicated professionals who lead the charge in digital workflows and enterprise solutions. As part of our team, you will engage in projects that merge technology with practical applications, creating revolutionary products that advance how services are delivered and managed.

Career Development

At ServiceNow, your career trajectory is filled with boundless opportunities. We support your growth with robust training programs, leadership development courses, and access to global challenges. Whether you are looking for an internship, full-time position, or leadership role, ServiceNow equips you with the tools to excel.

Be Part of a Great Team

Working at ServiceNow means being part of a community that values teamwork and innovation. Our collaborative environment encourages networking and sharing ideas, making our workplace vibrant and dynamic. The benefits of joining ServiceNow extend beyond comprehensive health and wellness; they include fostering professional connections and friendships that last a lifetime.

Explore Job Opportunities and Internships

Whether you’re a seasoned professional or a recent graduate, ServiceNow offers a range of employment options to suit your career goals. From internships that provide real-world experience to full-time positions that challenge you to leverage your expertise, we are committed to hiring the best talent.

Stay Connected

Join Our Team Search open positions that match your skills and interests. At ServiceNow, we look for passionate, curious, and solution-driven team players. Explore the possibilities that await you at a company that is committed to your professional success.

Keep Up to Date

Stay ahead with career tips, insider perspectives, and industry-leading insights you can put to use today—all from the people who work here.

Job Alert Emails

Customize your subscription to receive job alerts, the latest news, and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities that await at ServiceNow.

ServiceNow Careers

Empowering professionals to achieve more, ServiceNow is where careers are future-proofed, and ambitions are realized. Join us in our journey of growth and innovation.
Learn more about ServiceNow
Size
16,881 employees
Market Cap
$76.5 billion
Industry
Net Income
$118.5 million
Founded
2004
5 Year Trend
+33.5%
Revenue
$4.5 billion
NASDAQ

Similar Jobs

More Jobs at ServiceNow

More Information Technology Jobs

Find similar Manager, Network Reliability and Resiliency jobs: