Role Overview:Seeking an Observability Operations Engineer to administer and optimize enterprise observability platforms including Dynatrace, Splunk, and OpenSearch/Elasticsearch. This role involves ensuring platform availability, scalability, and security, building comprehensive monitoring solutions, and troubleshooting production issues in Linux, Kubernetes, and cloud environments.
Key Responsibilities:- Administer and optimize enterprise Dynatrace, Splunk, and OpenSearch/Elasticsearch platforms.
- Maintain platform availability, scalability, performance, security, and reliability.
- Build and manage monitoring, logging, tracing, dashboards, alerts, and operational metrics.
- Troubleshoot production issues and perform root cause analysis using observability tools.
- Support Linux, Kubernetes, container, and cloud-based environments.
- Automate operational activities and drive self-healing and AI-assisted operations.
- Manage upgrades, patching, capacity planning, backups, and operational governance.
- Collaborate with SRE, DevOps, Platform, Infrastructure, and Application teams.
Required Skills:- Strong Observability Administration experience with Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Hands-on experience with monitoring, logging, tracing, alerting, dashboards, and platform performance tuning.
- Strong knowledge of Linux, Kubernetes, and cloud environments.
- Experience with Grafana, Prometheus, OpenTelemetry, and related observability technologies.
- Automation experience using Python/Shell scripting and REST APIs.
- Experience supporting enterprise-scale production environments, troubleshooting, and RCA.
Qualifications:- 7 to 10 years of relevant experience.
Preferred Skills:- Strong Observability Administration experience with Dynatrace, Splunk, and OpenSearch/Elasticsearch.
- Hands-on experience with monitoring, logging, tracing, alerting, dashboards, and platform performance tuning.
- Strong knowledge of Linux, Kubernetes, and cloud environments.
- Experience with Grafana, Prometheus, OpenTelemetry, and related observability technologies.
- Automation experience using Python/Shell scripting and REST APIs.
- Experience supporting enterprise-scale production environments, troubleshooting, and RCA.