Netflix

Software Engineer L5 - AI Observability & Agent Evaluation

Netflix$388K — $500K+*
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software, AI/ML, or platform engineering.
  • Proven experience in designing standards, libraries, SDKs, or frameworks.
  • Strong coding skills in Python and experience with Java, Go, or Scala.
  • Practical experience with production ML models and data quality management.
  • Experience with observability stacks like Prometheus or Datadog.
  • Solid understanding of distributed systems and cloud platforms (AWS, GCP, Azure).
  • Ability to work cross-functionally and communicate complex system behaviors.

Responsibilities

  • Build observability framework for consistent metrics and logs in ML and GenAI systems.
  • Create primitives for teams to independently monitor model performance and data quality.
  • Develop evaluation frameworks for LLMs supporting quality and correctness assessments.
  • Lead evaluations for observability tools and integrate vendor platforms via SDKs and APIs.
  • Create reusable libraries and templates to promote observability as a default.
  • Provide dashboards, alerts, and templates for tracking model performance and reliability.

Benefits

  • Comprehensive health plans and mental health support.
  • 401(k) retirement plan with employer match.
  • Stock option program and disability benefits.
  • Health savings and flexible spending accounts.
  • Family-forming benefits and life insurance offerings.
  • Generous paid leave and flexible time off policies.
Full Job Description

The Opportunity

The AI Observability team makes AI, ML, and Agentic systems transparent, reliable, and production-ready at scale. We build end-to-end observability for ML and GenAI workloads, capturing model inputs, features, predictions, outcomes, and behavior across online and batch systems. Our platform enables teams to monitor model performance, data quality, drift, latency, and failures, turning the ML system from a black box into an explainable, debuggable system. We provide developer-friendly libraries, dashboards, and alerts so teams can debug issues, respond to incidents, and ship AI-powered products with confidence.

We're looking for a hands-on senior engineer to build the frameworks behind Netflix's AI Observability platform, model performance, evaluation, and vendor integration surfaces. You will design reusable infrastructure that enables ML/AI practitioners across domains to monitor model quality in production, evaluate LLM and agentic systems, and adopt vendor tooling through a consistent, self-serve platform. AIP owns the generic, reusable infrastructure; domain teams own their domain-specific evals and remediation. You will partner closely with engineering, product, machine learning, and data teams to turn their needs into reusable platform capabilities. To succeed, you will bring a strong background in AI or ML infrastructure and a passion for building scalable, robust systems.


In this role, you will:
  • Build the observability framework and platform capabilities that give ML and GenAI systems metrics, logs, and distributed traces across online inference, batch scoring, feature pipelines, and agent orchestration, so teams can instrument their systems consistently.

  • Build the primitives that let teams monitor model performance (accuracy, calibration, error rates), data quality, drift, and degradation on their own systems, rather than monitoring individual models yourself.

  • Build and extend evaluation frameworks for LLM and agentic systems that support response quality, grounding and hallucination, task success, tool-use and trajectory correctness, and LLM-as-a-judge and human-in-the-loop scoring, giving teams reusable building blocks to define and run their own evals.

  • Lead build-vs-buy evaluations for observability and eval tooling, and own the SDKs, connectors, and APIs that integrate vendor platforms into a consistent, well-supported interface, so ML and product teams can onboard models and agents with minimal friction.

  • Build reusable libraries, SDKs, and templates that make observability and evaluation the default for new systems (4observability-by-default4), lowering the barrier for teams to instrument and evaluate their work.

  • Provide the dashboarding, alerting, and SLO/SLI building blocks (plus sensible out-of-the-box templates) that teams use to track model performance, latency, cost, and reliability.

To succeed in this role, you will need:
  • Experience in software, AI/ML, or platform engineering, with hands-on time in production observability, monitoring, or ML/LLM evaluation

  • Proven track record of designing standards, libraries, SDKs, or frameworks that other teams build on and adopt, reflecting a platform and enablement mindset rather than shipping features for a single use case.

  • Strong coding skills in Python and at least one of Java, Go, or Scala, with experience building production services

  • Practical experience operating ML models in production (online serving and/or batch), including drift, model performance, and data quality.

  • Hands-on experience with modern observability stacks (e.g., Prometheus/Grafana, Datadog, OpenTelemetry, ELK/OpenSearch, Jaeger/Tempo, or similar).

  • Solid understanding of distributed systems, microservices, and at least one major cloud platform (AWS, GCP, or Azure).

  • Ability to work cross-functionally with ML, data, infra, and product teams, and to communicate clearly about system behavior, quality, and risk.

  • Experience with vendor integration and VPC deployment.

  • AI-Native Engineering Mindset who uses AI tools as a core part of their own workflow to accelerate design, development, testing, and code review.


Nice to have:
  • Hands-on experience with ML/LLM observability and evaluation tools (e.g., Arize, Braintrust, LangFuse, Weights & Biases, Galileo, Vertex AI Model Monitoring, SageMaker Model Monitor).

  • Experience building or shipping LLM/GenAI applications and evaluating them: prompt/result logging, evaluation metrics, LLM-as-a-judge, and human-in-the-loop review.

  • Experience evaluating agentic systems: tool use, multi-step reasoning, and trajectory/task-success measurement.

To learn more about our AI Platform, you can review the relevant talks/blog posts on the .

Generally, our compensation structure consists solely of an annual salary; we do not have bonuses. You choose each year how much of your compensation you want in salary versus stock options. To determine your personal top of market compensation, we rely on market indicators and consider your specific job family, background, skills, and experience to determine your compensation in the market range. The range for this role is $388,000.00 - $619,000.00.

Netflix provides comprehensive benefits including Health Plans, Mental Health support, a 401(k) Retirement Plan with employer match, Stock Option Program, Disability Programs, Health Savings and Flexible Spending Accounts, Family-forming benefits, and Life and Serious Injury Benefits. We also offer paid leave of absence programs. Full-time hourly employees accrue 35 days annually for paid time off to be used for vacation, holidays, and sick paid time off. Full-time salaried employees are immediately entitled to flexible time off. See more details about our Benefits here.

About Netflix

Netflix, Inc. is an American media company founded on August 29, 1997 by Reed Hastings and Marc Randolph in Scotts Valley, California, and currently based in Los Gatos, California, with production offices and stages at the Los Angeles-based Hollywood studios (formerly old Warner Brothers studios) and the Albuquerque Studios (formerly ABQ studios). It operates an eponymous over-the-top subscription video on-demand service, which showcases acquired and original programming as well as third-party content licensed from other production companies and distributors. Netflix is also the first streaming media company to be a member of the Motion Picture Association.
Learn more about Netflix
Size
11,300 employees
Market Cap
$127.6 billion
Industry
Net Income
$2.7 billion
Founded
1997
5 Year Trend
+27.5%
Revenue
$24.9 billion
NASDAQ

Similar Jobs

More Jobs at Netflix

More Information Technology Jobs

Find similar Software Engineer L5 - AI Observability & Agent Evaluation jobs: