AIOps & Observability Lead

Bessemer Trust

• $140K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of hands-on experience with observability and event management tools (e.g., Splunk, BigPanda, PagerDuty).
  • Experience designing and implementing AIOps capabilities for better alert correlation and automation.
  • Proficient in defining incident management processes with a focus on major incident response.
  • Ability to operate both strategically and technically in tooling settings.
  • Experience in scripting or automation (Python, PowerShell) for event enrichment and automated remediation.
  • Strong communication skills to align teams around an observability strategy.

Responsibilities

  • Define and oversee the observability and event management strategy.
  • Lead the rollout of Splunk Observability across applications and infrastructure.
  • Implement event management tools for alert correlation and prioritization.
  • Coach the Automation team on observability and event intelligence platforms.
  • Collaborate with the Infrastructure Architect to develop a full-scale observability roadmap.
  • Create and execute an AIOps strategy for automated, self-healing operations.
  • Manage partnerships with service providers for effective monitoring and escalation.
  • Design escalation paths to optimize L3 engineer time.
  • Establish standards for event classification and management across systems.
  • Explore expansion into desktop observability with Nexthink.

Benefits

  • 401(k) program with generous profit-sharing contribution.
  • Comprehensive medical, dental, and vision coverage.
  • Life and disability insurance coverage.
  • Paid holidays, vacation, and sick time.
Full Job Description
Role Summary

We are building a new Operations function and need a hands-on, technical leader to set the direction for our AIOps and observability strategy. This is a player-coach role, weighted toward player: someone who works directly in the tools, not just the roadmap.

Our new Splunk Observability implementation is a foundation of a broader automated operational intelligence ecosystem. We will lean on managed-service partners to act as first-line "eyes on glass" for the alerts and events it generates - but this role owns the strategy, architecture, and quality behind what those partners are watching. The goal is not just detection; it's ensuring issues can be discovered, triaged, and resolved through tooling and automation, reducing L3 engineer escalations wherever possible.

The role could expand over time into desktop/endpoint observability, including a Nexthink implementation.

Key Responsibilities

  • Define and own the observability and event management strategy, stack, and processes, including how major incidents are detected, triaged, and resolved.
  • Lead the rollout of Splunk Observability across our application and infrastructure portfolio, driving onboarding, instrumentation, and coverage expansion.
  • Lead implementation of event management tooling - such as, BigPanda, Splunk ITSI, or PagerDuty - to correlate, prioritize, and route alerts from Splunk Observability.
  • Coach and enable the Automation team to become proficient in the observability and event intelligence platform, including alert logic, dashboard creation, event enrichment, correlation workflows, triggered automations, MCP-style integrations, and runbook automation.
  • Partner with the Infrastructure Architect to build a roadmap toward first-class, enterprise-grade observability.
  • Develop and execute an AIOps strategy that turns telemetry into HITL-automated, self-healing operations - evaluating and building on tooling in Splunk, AWS, or elsewhere as appropriate.
  • Stand up and mature an AIOps/SRE capability: automated remediation, intelligent alert correlation, and reduced manual triage.
  • Direct managed-service partners providing "eyes on glass" monitoring - defining what they watch, how they escalate, and holding them accountable to SLAs.
  • Design escalation paths so that issues resolvable via tooling are handled there first, protecting L3 engineering time for what truly needs it.
  • Establish standards for event classification, severity, ownership, suppression, and closure across the environment.
  • Assess and help scope the expansion into desktop/endpoint observability, including a Nexthink implementation.
  • Ensure dashboards, alerts, and event workflows are built around operational decisions, not just technical visibility.
  • Align observability, event management, incident, problem, change, knowledge, and escalation workflows with ITIL/ITSM practices and ServiceNow processes.

Required Qualifications

  • Strong, hands-on experience with observability, monitoring, and event management platforms (e.g., Splunk, Splunk ITSI, BigPanda, PagerDuty).
  • Proven experience designing and implementing AIOps or event-intelligence capabilities - alert correlation, noise reduction, and automation.
  • Experience defining incident management processes, including major incident response and escalation design.
  • Comfortable operating as both strategic lead and hands-on technical contributor - able to shape direction and work directly in the tools.
  • Scripting or automation experience (e.g., Python, PowerShell, REST APIs) to support event enrichment and automated remediation.
  • Strong communication skills, with the ability to align infrastructure, application, and operations teams around a shared observability strategy.

Preferred Qualifications

  • Experience with AWS or other cloud-native observability and automation tooling.
  • Experience with endpoint/desktop observability platforms such as Nexthink.
  • ITIL/ITSM familiarity, particularly incident, problem, and change management.
  • Experience managing or directing managed-service/outsourced monitoring partners.
  • Experience in regulated, high-availability, or large-scale enterprise environments.

Success Profile

The right person has been in the room during major incidents, knows what a noisy, low-trust alerting environment looks like, and knows what it takes to fix it. They can set a multi-year observability direction and also configure the tool themselves when needed. Success looks like fewer incidents reaching L3, faster time-to-resolution, and an Operations function that trusts its own data.

The base salary range for this position is $140,000 - $180,000 per year. This range reflects the minimum and maximum base salary we reasonably expect to pay for this role. In addition, this position may be eligible to participate in the relevant business unit's incentive compensation plan, and other compensation programs as applicable. Eligible employees may participate in a 401(k) program with a generous profit-sharing contribution, medical, prescription dental, and vision coverage; life insurance; disability coverage; paid holidays; vacation; and sick time, subject to plan terms and Company policies.

Similar Jobs

More Jobs at Bessemer Trust

More Information Technology Jobs

Find similar AIOps & Observability Lead jobs: