Kyndryl Holdings, Inc.

Associate Director, Observability and Service Reliability

Kyndryl Holdings, Inc.$125K — $150K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in IT, Computer Science, Engineering, or equivalent experience.
  • 8+ years of experience in observability, application performance, or enterprise technology operations.
  • 3+ years in technical or people leadership roles.
  • Experience designing monitoring capabilities in complex environments.
  • Proficiency in Dynatrace and familiarity with Nexthink or similar technologies.
  • Understanding of observability architectures and infrastructure monitoring.
  • Ability to influence senior leaders and articulate risk and outcomes.

Responsibilities

  • Own and evolve the enterprise observability and service reliability strategy.
  • Define monitoring standards for critical technology services and platforms.
  • Identify gaps and opportunities to enhance observability across the organization.
  • Establish a service reliability framework and work with teams on monitoring needs.
  • Lead the strategy and optimization for observability tools like Dynatrace and Nexthink.
  • Drive proactive incident response and automation of monitoring processes.
  • Create governance around monitoring standards and support continuous improvement.

Benefits

  • Opportunity to work with cutting-edge observability and monitoring technologies.
  • Impactful role in shaping the enterprise's observability and reliability strategy.
  • Chance to influence technical design across various teams and services.
  • Collaboration with diverse teams in application, infrastructure, and cybersecurity.
  • Professional growth in a leadership capacity with a focus on automation and analytics.
Full Job Description
The Role
Enterprise Observability Strategy
  • Own and mature the enterprise observability and service reliability strategy.
  • Define standards for monitoring applications, infrastructure, cloud platforms, networks, endpoints, APIs, databases, middleware, and other critical technology services.
  • Establish expectations for metrics, logs, traces, events, synthetic monitoring, real user monitoring, digital experience, service health, and business transaction visibility.
  • Create a consistent enterprise approach while allowing teams to use monitoring technologies suited to their platforms and services.
  • Identify monitoring gaps, redundant capabilities, excessive alerting, and opportunities to improve visibility.
  • Move the organization from traditional monitoring toward proactive, predictive, and automated operations.
Service Reliability Engineering
  • Establish and mature the organization's service reliability framework.
  • Partner with technical service owners to define monitoring requirements for critical applications and services.
  • Ensure monitoring reflects the complete service, including application performance, infrastructure, dependencies, integrations, user experience, business transactions, capacity, and failure conditions.
  • Define minimum observability requirements based on service criticality and business impact.
  • Help teams establish meaningful Service Level Indicators, Service Level Objectives, availability targets, performance thresholds, and health measures.
  • Use reliability data, incidents, problem records, capacity trends, and telemetry to identify systemic weaknesses and prioritize improvements.
Technical Thought Leadership
  • Serve as the enterprise technical authority for observability, monitoring, and service reliability.
  • Provide architectural guidance to application, infrastructure, cloud, engineering, DevOps, SRE, cybersecurity, and operations teams.
  • Influence solution design so services are observable, measurable, supportable, and resilient by design.
  • Develop enterprise monitoring patterns, reference architectures, standards, and reusable capabilities.
  • Guide technical teams in selecting appropriate monitoring methods and technologies for specific platforms and use cases.
  • Evaluate emerging observability, AIOps, automation, analytics, and service reliability capabilities for measurable operational value.
Observability Platform Leadership
  • Provide strategic oversight for the enterprise observability and monitoring tool ecosystem.
  • Lead the strategy, architecture, governance, adoption, and optimization of major platforms, including Dynatrace, Nexthink, and related enterprise monitoring technologies.
  • Ensure monitoring tools operate as an integrated ecosystem rather than isolated platforms.
  • Establish standards for instrumentation, tagging, alerting, dashboards, integrations, service mapping, ownership, and data quality.
  • Partner with technical teams to maximize platform value while reducing tooling duplication and complexity.
  • Manage strategic technology and vendor relationships to ensure observability investments deliver measurable operational value.
Dynatrace Platform Strategy
  • Provide strategic leadership for enterprise use of Dynatrace across applications, infrastructure, cloud, and digital services.
  • Drive adoption of application performance monitoring, distributed tracing, real user monitoring, synthetic monitoring, infrastructure monitoring, logs, topology, service health, and intelligent problem detection.
  • Partner with technical owners to improve application instrumentation and ensure Dynatrace provides meaningful service visibility.
  • Use Dynatrace capabilities to improve root cause identification, dependency visibility, anomaly detection, and incident response.
Nexthink and Digital Employee Experience
  • Lead the enterprise strategy for Nexthink and Digital Employee Experience.
  • Provide visibility into endpoint health, application performance, employee technology experience, device reliability, and technology friction.
  • Partner with Digital Workplace, Service Desk, application teams, and infrastructure teams to identify and remediate issues proactively.
  • Use experience data to reduce support demand, improve employee productivity, and identify systemic technology issues.
  • Expand automation and targeted remediation to resolve employee experience issues at scale.
Event Management, AIOps, and Automation
  • Improve operational signal quality by reducing noise and increasing alert relevance.
  • Drive event correlation, enrichment, anomaly detection, automated diagnostics, and proactive remediation.
  • Integrate observability platforms with ITSM, incident management, automation, collaboration, configuration, and operational data platforms.
  • Develop capabilities that help support teams detect degradation earlier and understand business impact faster.
  • Use AI, machine learning, analytics, and automation where they improve outcomes and reduce manual effort.
Governance, Metrics, and Continuous Improvement
  • Establish governance for enterprise monitoring standards, tooling, integrations, data quality, licensing, and adoption.
  • Develop executive and operational reporting that provides insight into service health, reliability, performance, and user experience.
  • Use service measures to identify reliability risks, guide investment decisions, and prioritize service improvements.

Key measures may include:
  • Service availability and reliability
  • Mean Time to Detect, Mean Time to Identify, and Mean Time to Restore
  • Monitoring coverage for critical services
  • Service Level Objective attainment
  • Alert quality and noise reduction
  • Proactive issue detection and automated remediation
  • Digital Employee Experience
  • Recurring service degradation
  • Observability maturity
  • Tool adoption and value realization
Leadership Responsibilities
  • Lead and develop observability, monitoring, reliability, and platform engineering professionals.
  • Create an engineering culture focused on proactive operations, technical excellence, automation, and measurable service outcomes.
  • Build strong partnerships with application owners, platform owners, infrastructure, cloud, network, cybersecurity, Digital Workplace, DevOps, SRE, Service Management, and business teams.
  • Influence technical teams without relying solely on direct reporting authority.
  • Create clear accountability for monitoring coverage and service reliability across the enterprise.
  • Partner with Incident, Problem, Change, Major Incident Management, and Operational Resilience teams to improve detection, recovery, and prevention.


Who You Are
Required Qualifications
  • Bachelor's degree in Information Technology, Computer Science, Engineering, or a related discipline, or equivalent professional experience.
  • 8 or more years of experience in observability, application performance, service reliability, infrastructure, cloud, engineering, or enterprise technology operations.
  • 3 or more years of technical leadership or people leadership experience.
  • Experience designing and operating monitoring and observability capabilities in complex enterprise environments.
  • Experience with Dynatrace and familiarity with Nexthink or comparable Digital Employee Experience technologies.
  • Knowledge of observability technologies, monitoring architectures, application performance management, infrastructure monitoring, event management, logging, tracing, synthetic monitoring, and user experience monitoring.
  • Understanding of modern enterprise applications, cloud technologies, containers, APIs, networks, databases, infrastructure, and distributed architectures.
  • Experience defining service health, reliability measures, Service Level Indicators, Service Level Objectives, and operational performance standards.
  • Ability to influence senior technical leaders and translate complex technical concepts into business risk and operational outcomes.
  • Strong leadership, architecture, analytical, problem-solving, communication, and stakeholder management skills.
Preferred Qualifications
  • Advanced Dynatrace experience in a large enterprise environment.
  • Experience implementing or scaling Nexthink.
  • Experience with ServiceNow and enterprise event management platforms.
  • Experience with Site Reliability Engineering, DevOps, AIOps, automation, OpenTelemetry, cloud native monitoring, and modern observability architectures.
  • Experience establishing enterprise monitoring standards or observability reference architectures.
  • Experience within complex, global, or highly regulated enterprise environments.

About Kyndryl Holdings, Inc.

Kyndryl is an IT transformation services company. They offer network services, business resilience, and hybrid cloud solutions. Their services include applications, core enterprise and zcloud, the digital workplace, security and resiliency, cloud, data, and AI, network, and edge. They committed to the health and continuous improvement of the vital systems at the heart of the digital economy.

Kyndryl Holdings, Inc. Careers

Join the Pioneering Team at Kyndryl Holdings, Inc. At Kyndryl Holdings, Inc., we are at the forefront of technological innovation and service excellence, making this an ideal time to join our global team. As a leader in IT infrastructure services, we offer unparalleled job opportunities that propel your career to new heights. Work You’ll Do Embark on a career with Kyndryl Holdings, Inc. and be part of a team that transforms businesses and industries worldwide. Our professionals are not just part of a company; they are essential drivers of innovation and growth. With a commitment to leadership and diversity training, Kyndryl Holdings, Inc. stands out as a beacon of professional development and innovation in the tech industry. Our culture is built on the principles of open collaboration, leadership, and respect for diversity, ensuring a welcoming environment for all employees. Join our team and contribute to our mission of delivering high-quality solutions across various industries. At Kyndryl Holdings, Inc., your skills will be honed through challenging projects and a collaborative team environment. Innovative Work At Kyndryl Holdings, Inc., innovation isn't just a buzzword—it's the cornerstone of everything we do. With a focus on cutting-edge technology and industry expertise, our team is equipped to tackle some of the most critical challenges facing businesses today. Be Part of a Great Team Kyndryl Holdings, Inc. is not just about individual growth; it’s about collective success. The synergy of our team’s diverse skills leads to groundbreaking solutions and industry leadership. Our commitment to employee benefits and a supportive culture makes Kyndryl Holdings, Inc. a top choice for professionals seeking meaningful and rewarding careers. Future-Proof Your Career With Kyndryl Holdings, Inc., the path to professional advancement is clear. Our career development programs offer extensive training, certification support, and opportunities for growth, ensuring that your career trajectory is always upward. Explore Job Opportunities Whether you’re looking for a full-time position, an internship, or leadership roles, Kyndryl Holdings, Inc. offers a range of employment opportunities to match your career aspirations. Enhance your resume through hands-on experience that sets you apart in the job market. Kyndryl Holdings, Inc. Careers Network Stay connected with Kyndryl Holdings, Inc. through our dedicated careers network. Search open positions that align with your skills and interests. We are continuously hiring creative, curious, and motivated individuals who are ready to drive success in a dynamic work environment. Keep Up to Date Stay informed with the latest career tips, industry insights, and professional growth opportunities available at Kyndryl Holdings, Inc. Our careers blog provides valuable information to help you navigate your job search, prepare for interviews, and excel in your career. Job Alert Emails Customize your experience by subscribing to job alerts and insider tips tailored to your preferences. Discover the exciting and rewarding career opportunities waiting for you at Kyndryl Holdings, Inc. Join Kyndryl Holdings, Inc. today and be part of a team that values innovation, leadership, and a diverse skill set in driving technological advancements and business success.
Learn more about Kyndryl Holdings, Inc.
Market Cap
$2.5 billion
Industry
NASDAQ

Similar Jobs

More Jobs at Kyndryl Holdings, Inc.

More Information Technology Jobs

Find similar Associate Director, Observability and Service Reliability jobs: