Application Monitoring Engineer

D2 Technical Services

$150K — $175K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in application monitoring, observability, or site reliability engineering
  • Hands-on expertise in major APM tools (Datadog, New Relic, Dynatrace)
  • Experience with cloud platforms (AWS, Azure, GCP) and containerization (Docker, Kubernetes)
  • Proficient in scripting or programming (Python, Bash) for automation and integration
  • Solid understanding of distributed systems and microservices architecture
  • Experience in on-call rotations and incident response
  • Ability to communicate complex technical issues clearly to varying audiences

Responsibilities

  • Design and maintain end-to-end application performance monitoring
  • Define and manage SLIs, SLOs, and error budgets in collaboration with teams
  • Lead triage and root-cause analysis for incidents using APM data
  • Drive continuous improvement initiatives for the observability stack
  • Mentor engineers on integrating effective monitoring into services from inception
  • Collaborate with DevOps/SRE and development teams to ensure monitoring aligns with infrastructure changes

Benefits

  • Health, Dental, and Vision Insurance
  • 401(k) matching
  • Accrued Paid Time Off (PTO)
  • Short-Term and Long-Term Disability Insurance
  • Life Insurance
  • Referral Bonuses
  • Professional development reimbursement
Full Job Description
**ACTIVE TS/SCI SECURITY CLEARANCE REQUIRED**

We're looking for a Senior Application Monitoring Engineer to own and evolve our application observability strategy. You'll be the go-to expert for keeping critical business applications visible, healthy, and performant - designing monitoring architecture, driving down mean time to detection and resolution, and mentoring other engineers on observability best practices. This is a largely autonomous role for someone who thinks proactively about failure modes and treats monitoring as a product, not an afterthought.
What You'll Do
  • Design, build, and maintain end-to-end application performance monitoring using platforms like Dynatrace, instrumenting applications and services so that the right signals surface before customers notice a problem
  • Define and maintain SLIs, SLOs, and error budgets in partnership with engineering and product teams, and build dashboards and alerting that reduce noise while catching what actually matters
  • When incidents happen, you'll lead or support triage and root-cause analysis, using distributed tracing and APM data to pinpoint issues across complex, distributed systems.
  • Drive continuous improvement of the observability stack itself - evaluating new tools, refining alert thresholds, and reducing alert fatigue across the engineering organization
  • Serve as a technical mentor, helping other engineers build monitoring into their own services from the start rather than bolting it on later
  • Partner closely with DevOps/SRE, infrastructure, and application development teams to make sure monitoring coverage keeps pace with new deployments and architectural changes
What We're Looking For
  • 5+ years of experience in application monitoring, observability, or site reliability engineering, with hands-on expertise in at least one major APM platform (Datadog, New Relic, or Dynatrace) - deep familiarity with more than one is a strong plus
  • Experience configuring alerts and adjusting thresholds
  • Solid grasp of distributed systems, microservices architecture, and how to trace a request across services, containers, and cloud infrastructure
  • Experience with scripting or programming (Python, Bash, or similar) to automate monitoring configuration and build custom integrations
  • Experience with cloud platforms (AWS, Azure, or GCP) and containerized environments (Docker, Kubernetes) is expected
  • Understanding of the fundamentals of APIs, databases, and networking well enough to diagnose issues across the full stack
  • You've participated in on-call rotations and incident response processes, and ideally have exposure to complementary tools like Splunk, ELK, Grafana, or Prometheus
  • Communicate clearly under pressure, can explain a complex incident to both engineers and non-technical stakeholders, and take genuine ownership of the systems you monitor
Nice to Have
  • Familiarity with CI/CD pipelines and infrastructure-as-code (Terraform, CloudFormation, etc.) is a plus
  • Relevant certifications (Datadog, New Relic, AWS/Azure/GCP)


Additional Information
  • Compensation is unique to each candidate and relative to the skills and experience they bring to the position. The salary range for this position is typically $150-175k. This does not guarantee a specific salary as compensation is based upon multiple factors such as education, experience, certifications, and other requirements, and may fall outside of the above-stated range.
  • Highlights of our benefits include Health/Dental/Vision, 401(k) match, Accrued PTO, STD/LTD/Life Insurance, Referral Bonuses, professional development reimbursement, and more!

Similar Jobs

More Jobs at D2 Technical Services

More Information Technology Jobs

Find similar Application Monitoring Engineer jobs: