Head of Enterprise Monitoring and ObservabilityReports ToHead of Technology Operations
DepartmentEnterprise Technology
Position SummaryThe Head of Enterprise Monitoring and Observability is responsible for defining and executing the enterprise observability strategy, operating model, and platform roadmap that enable proactive, predictive, and resilient technology operations.
This leader owns the enterprise capabilities for monitoring, telemetry, event intelligence, operational analytics, and observability engineering, ensuring technology services are observable, measurable, actionable, and continuously improving.
The role serves as the operational intelligence leader across Enterprise Technology, providing the platforms, standards, insights, and governance required to improve service reliability, accelerate issue detection and resolution, and support modern operational practices.
This leader partners closely with infrastructure, cloud, middleware, application, security, resiliency, and service management teams to build a unified enterprise-wide observability capability.
Key ResponsibilitiesEnterprise Observability Strategy- Define and execute the enterprise observability vision, strategy, operating model, and multi-year roadmap.
- Establish standards and governance for monitoring, logging, metrics, tracing, telemetry, alerting, and service observability.
- Lead the evolution of observability capabilities from reactive monitoring toward proactive, predictive, and automated operations.
- Drive enterprise adoption of modern observability practices, frameworks, and engineering standards.
Observability Platform Ownership- Own the strategy, architecture, lifecycle, and operation of enterprise monitoring and observability platforms.
- Lead platform modernization, tool rationalization, capability expansion, and service offering improvements.
- Manage vendor relationships, contracts, platform investments, and cost optimization efforts.
- Ensure scalability, availability, resilience, and ongoing enhancement of observability tooling and services.
Monitoring and Telemetry Services- Establish enterprise monitoring and instrumentation standards across infrastructure, cloud, middleware, applications, databases, and digital services.
- Improve monitoring coverage, telemetry quality, alert effectiveness, signal quality, and operational visibility.
- Drive adoption of distributed tracing, service dependency mapping, and end-to-end transaction monitoring.
- Ensure service health and operational data are consistently available across the technology landscape.
Operational Intelligence and AIOps- Lead implementation of operational analytics, event intelligence, anomaly detection, predictive insights, and AIOps capabilities.
- Develop dashboards, scorecards, and reporting that provide actionable operational intelligence to technology leadership.
- Enable intelligent event correlation, noise reduction, automated diagnostics, and automated operational workflows.
- Partner with engineering and operations teams to use observability data for continuous improvement and outage prevention.
Reliability and Incident Enablement- Partner with technology teams to improve detection, response, recovery, and root cause analysis capabilities.
- Establish observability practices that improve service reliability, operational resilience, and customer experience.
- Drive post-incident learning and identification of automation opportunities.
- Provide enterprise visibility into technology health, performance trends, capacity risks, and reliability risks.
Leadership and Talent Development- Build, lead, and develop a high-performing team of observability engineers, platform engineers, and operational intelligence specialists.
- Foster a culture of innovation, accountability, operational excellence, continuous learning, and cross-functional partnership.
- Lead organizational adoption of observability best practices through enablement, coaching, and stakeholder engagement.
- Establish strong partnerships across technology, cybersecurity, architecture, engineering, operations, resiliency, and service management functions.
Qualifications- 10+ years of experience in infrastructure operations, platform engineering, observability, reliability engineering, or related technology disciplines.
- 5+ years of experience leading enterprise-scale engineering or operational teams.
- Deep expertise in observability platforms, monitoring technologies, telemetry pipelines, logging, metrics, distributed tracing, and operational analytics.
- Demonstrated experience implementing enterprise monitoring strategies and improving operational reliability at scale.
- Strong understanding of cloud platforms, hybrid infrastructure, automation, DevOps practices, and modern operational architectures.
- Experience leading organizational transformation and driving adoption of new operational capabilities.
- Excellent leadership, communication, stakeholder management, and strategic planning skills.
Success Measures- Increased enterprise monitoring and observability coverage.
- Reduced mean time to detect (MTTD) and mean time to resolve (MTTR).
- Improved signal quality and reduced alert fatigue.
- Increased adoption of observability standards and platform capabilities.
- Improved service reliability, operational resilience, and operational efficiency.
- Expansion of predictive analytics, automation, and AIOps capabilities across Enterprise Technology.
Salary Range:$110,350.00 - $181,285.00
The salary range reflected above is a good faith estimate of base pay for the primary location of the position. The salary for this position ultimately will be determined based on the education, experience, knowledge, and abilities of the successful candidate. In addition to salary, this role may also be eligible for annual, sales, or other incentive compensation.