We are seeking a Mid-Level Observability Engineer to help build, maintain, and enhance our enterprise monitoring and observability capabilities across cloud and hybrid environments. This role is hands-on and execution-focused, supporting Dynatrace and AWS CloudWatch implementations, dashboard development, alert tuning, instrumentation, and operational reporting for critical platforms and applications. The ideal candidate will partner with Cloud Engineering, DevOps, SRE, and application teams to improve service visibility, strengthen monitoring coverage, and embed observability practices into ongoing operational and delivery workflows.
Observability Platform Ownership & Architecture
- Design, deploy, and optimize enterprise-grade observability solutions using Dynatrace SaaS or Managed, including OneAgent, ActiveGate, full-stack monitoring, RUM, synthetic monitoring, Davis AI, distributed tracing, dashboards, and log monitoring on Grail.
- Define platform standards for tagging, management zones, network segmentation, alerting profiles, access control, dashboards, and telemetry governance across hybrid environments.
- Architect observability coverage across AWS and on-prem platforms, including containerized and serverless workloads such as EKS, ECS, Lambda, EC2, RDS, and API Gateway.
- Lead migration from legacy monitoring tools into Dynatrace and drive closure of enterprise monitoring gaps through structured onboarding and platform modernization. 2. Application Performance Management & Incident Triage
- Configure and optimize APM instrumentation for distributed applications, APIs, microservices, databases, and business transactions.
- Serve as the escalation point for complex incidents, using Smartscape, Distributed Traces, Davis AI, Live Debugger, and method-level diagnostics to accelerate root cause identification and reduce MTTR.
- Define and maintain SLIs, SLOs, and error budgets, aligning platform telemetry to business reliability targets and engineering commitments
- Lead post-incident reviews using observability evidence and drive corrective improvements in instrumentation, thresholds, dashboards, and alerting logic3. Telemetry Automation & Observability as Code
- Standardize monitoring configurations using Terraform and/or Dynatrace Monaco, including alerting profiles, dashboards, SLOs, tagging rules, synthetic tests, and management zones
- Build automation for platform operations, integration workflows, reporting, and remediation using Python, Bash, or PowerShell, along with REST APIs and webhooks.
- 5–10 years of experience in Observability, Monitoring Engineering, SRE, APM, DevOps, or Infrastructure Engineering, including several years of hands-on Dynatrace administration and architecture.
- Deep hands-on expertise with Dynatrace across full-stack monitoring, Davis AI, Smartscape, RUM, synthetic monitoring, distributed tracing, Grail log monitoring, DQL, management zones, Workflows/AutomationEngine, and access governance.
- Strong experience with AWS cloud services, especially CloudWatch, EKS, ECS, Lambda, EC2, RDS, API Gateway, networking, and modern cloud architecture patterns.
- Advanced knowledge of Kubernetes and cloud-native observability patterns, including instrumentation for microservices and distributed systems.
- Strong proficiency in observability-as-code using Terraform and/or Monaco, plus scripting in Python, Bash, or PowerShell
- Solid understanding of distributed application architecture, networking fundamentals, telemetry pipelines, performance engineering, and incident management.
- Dynatrace certification at Associate or Professional level required; higher-level certification is strongly preferred.
Preferred Qualifications
- Experience with tools such as Nagios/SolarWinds/Prometheus/Grafana, Splunk, or ELK
- Experience with OpenTelemetry, Dynatrace Grail, advanced log analytics, and enterprise telemetry standardization
- Experience integrating observability with ITSM or event-management platforms such as ServiceNow
- Background in SRE practices such as reliability reviews, error budget management, and incident reduction programs.
- Integrate observability controls into CI/CD pipelines and establish telemetry quality standards for new application and infrastructure deployments
- Use Dynatrace Query Language (DQL) and Grail capabilities for advanced log analysis, event correlation, notebooks, and custom operational insights. 4. Governance, Cost Control & Enablement
- Own monitoring governance practices related to telemetry quality, alert design, data retention, platform usage standards, and operational reporting
- Manage Dynatrace usage and consumption responsibly by monitoring ingest patterns, tuning retention, and optimizing log, metric, and trace collection for value and efficiency.
- Build executive and engineering dashboards that communicate service health, reliability KPIs, error budgets, and infrastructure visibility to multiple audiences
- Mentor engineers and partner teams on observability best practices, onboarding, dashboarding, instrumentation, and platform self-sufficiency.
This opportunity is for an existing vacancy with the company. The anticipated base salary range for this position is$85,000 to $125,000. Exact salary depends on several factors such as experience, skills, education, and budget. Salary range may vary based on geographic location. In addition to base salary, this position is eligible for participation in a bonus program. In addition, The Company offers a variety of benefits to eligible employees, including health insurance coverage, wellness programs, life and disability insurance, retirement savings plans, paid leave programs, education-related programs, paid holidays and vacation time, and many others. Many of these benefits are subsidized or fully paid for by the company.