Top Skills Required:1.DataDog
2.Splunk
3. AIOps Platforms
Role OverviewWe are seeking an experienced Enterprise Observability Architect to lead enterprise observability initiatives across infrastructure, applications, cloud, and platform services. The candidate will be responsible for defining observability architecture, establishing monitoring standards, implementing Datadog solutions, driving platform governance, and enabling AI-driven operational insights. The role requires strong consulting, architecture, and stakeholder management skills with the ability to lead large-scale observability transformation programs.
Key Responsibilities:- Define and own enterprise observability architecture, standards, and governance.
- Lead monitoring and observability assessments to identify gaps and improvement opportunities.
- Design and implement scalable Datadog monitoring solutions across infrastructure, applications, databases, middleware, networks, storage, backup, cloud, and collaboration platforms.
- Develop monitoring, alerting, logging, dashboarding, and reporting frameworks.
- Architect and implement Datadog capabilities including Infrastructure Monitoring, APM, Logs, Synthetic Monitoring, RUM, NPM, DBM, and Distributed Tracing.
- Drive observability adoption across hybrid cloud and enterprise platforms.
- Integrate observability solutions with ITSM, AIOps, automation, and AI/GenAI platforms.
- Develop predictive monitoring, anomaly detection, and operational intelligence use cases.
- Optimize monitoring coverage, alert quality, platform performance, and observability costs.
- Provide technical leadership, mentoring, and best practices to engineering and operations teams.
Required Skills:Experience10+ years of experience in Enterprise Monitoring and Observability.
6-7+ years of hands-on Datadog implementation and architecture experience.
Experience leading enterprise observability transformation programs.
Technical SkillsDatadog Infrastructure Monitoring
Application Performance Monitoring (APM)
Log Management
Distributed Tracing
Database Monitoring (DBM)
Network Performance Monitoring (NPM)
Real User Monitoring (RUM)
Synthetic Monitoring
Service Catalog & Service Mapping
SLO/SLI Design and Implementation
OpenTelemetry
Cloud & Platform ExpertiseAWS and Azure Cloud
Kubernetes and Container Monitoring
Linux, Windows, and Enterprise Infrastructure Monitoring
Database, Middleware, Storage, Backup, and Network Monitoring
Integration & AutomationSplunk
ServiceNow
AIOps Platforms
Python, PowerShell, Terraform, or Ansible
API-based Integrations