Enterprise Service Reliability and Insights Lead

$100K — $130K *
US-AnywhereRemote in United States
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Active Secret clearance (Tier 3 background investigation)
  • Bachelor's degree in IT, Cybersecurity, Systems Engineering, or related field
  • 10+ years in service reliability or IT operations
  • Experience with enterprise monitoring systems
  • Familiar with defining SLAs, KPIs, and performance metrics

Responsibilities

  • Define and manage enterprise monitoring and observability strategy
  • Oversee monitoring tools, dashboards, and alert configurations
  • Establish alert thresholds and performance indicators
  • Ensure monitoring meets uptime, performance, and security standards
  • Collaborate with engineering teams for incident management
  • Deliver insights on system reliability and performance issues
  • Participate in incident response and root cause analysis sessions

Benefits

  • Fully remote position
  • Collaborative and dynamic work environment
  • Opportunity to enhance monitoring and observability processes
  • Engagement with cutting-edge IT and cybersecurity practices
  • Involvement with federal and DoD-aligned IT projects
Full Job Description
Overview

DecisionPointseeks a Senior Enterprise Service Reliability and Insights Lead to oversee enterprise-wide monitoring, observability, and operational intelligence for a large federal and DoD-aligned IT environment. This senior-level role defines the monitoring strategy, manages toolsets, develops dashboards,establishesalerting thresholds, and ensures service reliability through proactive detection and rapid incident identification.

The Enterprise Service Reliability and Insights Leadis responsible fordriving visibility into uptime, system performance, service health, and operational risks. This position partners closely with Tier 2 and Tier 3 engineering teams, cloud operations, cybersecurity, and service desk leadership to ensure monitoring aligns with mission needs, SLAs, and enterprise performanceobjectives.

This position is fully remote.

Duties & Responsibilities

The Enterprise Service Reliability and Insights Lead will:

  • Define, implement, and managethe enterprisemonitoring and observability strategy.
  • Oversee monitoring tools, dashboards, agents, log pipelines, and alerting configurations across all environments.
  • Establish alert thresholds, escalation criteria, and performance indicators that support proactive issue detection.
  • Ensure monitoring coverage aligns with uptime, performance, and security requirements.
  • Collaborate with Tier 2 and Tier 3 engineering teams on system health assessments, log analytics, and incident triage.
  • Lead efforts to correlate events acrossapplication, infrastructure,network, and security monitoring tools.
  • Deliver actionable insights on system reliability, capacity issues, performance bottlenecks, and incident trends.
  • Support SLA and KPI measurement, reporting, and compliance tracking.
  • Maintain monitoring documentation, dashboards, service health definitions, and alerting standards.
  • Partner with cloud, infrastructure, and cybersecurity teams to ensure observability supports mission and compliance needs.
  • Recommend improvements tomonitoringarchitectures, event correlation, and automation capabilities.
  • Participate in incident response activities, root cause analysis sessions, and readiness reviews.
  • Drive continuous improvement initiatives across reliability engineering and service monitoring.
Qualifications

Clearance Requirement

Must hold an active Secret clearance, supported by a Tier 3 background investigation.

Education (Required)

Bachelor27s degree in Information Technology, Cybersecurity, Systems Engineering, ora relatedtechnical field.

Experience (Required)

  • Minimum 10 yearsof experience in service reliability,monitoringengineering, IT operations, or systems engineering.
  • Experience designing or managing enterprise monitoring systems and dashboards.
  • Experience defining SLAs, KPIs, and operational performance measurements.
  • Experience collaborating with Tier 2 and Tier 3 teams for incident management and problem resolution.
  • Experience with log analysis, event correlation, and observability platforms.

Technical Knowledge (Required)

  • Strong understanding of monitoring and observability tools (metrics, logs, traces).
  • Knowledge of uptime, performance, and reliability engineering practices.
  • Familiarity with ITIL v4 processes for incident, problem, and change management.
  • Understanding ofalerting strategies, threshold design, and escalation workflows.
  • Knowledge of DoD or federal IT operational environments.

Technical Knowledge (Preferred)

  • Experience with cloud-native monitoring services and distributed systemsmonitoring.
  • Experience with APM tools, SIEM integrations, or event correlation engines.
  • Familiarity with automation scripting or analytics for monitoring enhancement.

Certifications

Required:

  • ITIL v4 Foundation
  • CompTIA Security+

Preferred:

  • Cloud monitoring certifications (AWS, Azure, or similar)
  • SRE or observability-related certifications

Skills

  • Strong analytical skills for interpreting system health and service reliability data.
  • Excellent communication and reporting skills for executive and technical audiences.
  • Ability to lead cross-functional coordination during performance events and incidents.
  • High attention to detail with strong documentation habits.
  • Ability to drive continuous improvement across monitoring, reliability, and availability functions.

Similar Jobs

More Jobs at

More Enterprise Technology Jobs

Find similar Enterprise Service Reliability and Insights Lead jobs: