Observability and Evaluation Engineer

NTT Data, Inc.

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
  • Strong hands-on experience with Python programming.
  • Familiarity with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
  • Understanding of LLM evaluation and AI quality assessment methods.
  • Experience in Agile teams and production support environments.

Responsibilities

  • Implement observability and telemetry for LLM-powered applications and services.
  • Build evaluation suites to assess agent behavior and performance metrics.
  • Develop operational dashboards, alerts, and reporting for readiness.
  • Collaborate with platform engineers to establish monitoring and evaluation needs.
  • Automate the collection of evidence for release readiness and governance.
  • Create runbooks for priority agent releases and support documentation.
  • Analyze production data to recommend reliability and performance enhancements.

Benefits

  • Opportunities for professional development in a cutting-edge field.
  • Collaborative work environment with cross-functional teams.
  • Flexibility in work arrangements to support work-life balance.
Full Job Description
Job Description:
Position Summary

The Overwatch Observability & Evaluation Engineer will build telemetry, tracing, dashboards, evaluation suites, alerts, service objectives, runbooks, and readiness evidence for Tachyon agent releases. This role ensures production AI systems can be monitored, evaluated, improved, and supported with clear operational visibility.

Key Responsibilities
  • Implement observability and telemetry for LLM-powered applications, agents, tools, and platform services.
  • Build evaluation suites for agent behavior, prompt quality, response quality, retrieval performance, latency, reliability, and safety signals.
  • Develop dashboards, alerts, traces, metrics, service objectives, and reporting for production readiness.
  • Work with platform engineers and Product Owners to define monitoring requirements and evaluation metrics.
  • Automate evidence collection for release readiness, operational reviews, and governance checkpoints.
  • Create runbooks and support documentation for priority agent releases.
  • Analyze production behavior and recommend improvements to reliability, performance, and quality.

Required Qualifications
  • 7+ years of engineering experience with observability, monitoring, test automation, platform operations, or AI/ML systems.
  • Strong hands-on Python experience.
  • Experience with dashboards, metrics, alerts, traces, logs, SLOs, and production monitoring.
  • Understanding of LLM evaluation, prompt evaluation, RAG evaluation, or AI quality assessment approaches.
  • Experience working in Agile engineering teams and production support environments.

Required Skills / Knowledge
  • Python, telemetry, tracing, monitoring, dashboards, alerting, SLOs, evaluation frameworks, test automation, and production operations.
  • Understanding of LLMs, agents, RAG, prompt performance, retrieval quality, latency, and reliability metrics.
  • Experience with observability tools and open telemetry concepts.

Preferred Qualifications
  • Experience with GenAI observability, AI evaluation tools, ML monitoring, or platform reliability engineering.
  • Experience in regulated environments with evidence and readiness documentation.
  • Kubernetes, cloud platforms, and CI/CD experience.

Expected Outcomes
  • Operational dashboards and evaluation suites for priority agent releases.
  • Clear readiness evidence, alerts, SLOs, and runbooks.
  • Improved quality, reliability, and trust in production Agentic AI systems.


Job Code

Similar Jobs

More Jobs at NTT Data, Inc.

More Information Technology Jobs

Find similar Observability and Evaluation Engineer jobs: