Nuna Incorporated

Software Engineer, AI Evaluation

Nuna Incorporated$140K — $170K *
Healthcare
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Significant experience in developing and deploying reliable production systems and tooling.
  • Expertise in evaluating AI systems with knowledge of adversarial testing and scenario generation.
  • Strong testing mindset focused on building measurement systems, not just running tests.
  • Proficiency in utilizing AI tools to enhance team effectiveness and workflows.
  • Fluency in statistics and experimental design to collaborate with data scientists.
  • Ability to design user-friendly workflows and interfaces for non-engineers.
  • Passion for advancing healthcare with a keen sense of critical issues versus nice-to-haves.

Responsibilities

  • Construct testing harnesses and evaluation infrastructure for AI products.
  • Oversee end-to-end evaluations, including architecture and content, collaborating with data scientists and clinicians.
  • Ensure all AI deployments undergo rigorous evaluation before release, managing safety and quality controls.
  • Develop ground truth metrics and validation processes, ensuring evaluation reliability and calibrating against human labels.
  • Create functional tools for clinicians and designers to author review scenarios without engineering support.
  • Facilitate the iterative improvements from evaluation results to model refinement, enhancing autonomy in systems.

Benefits

  • Opportunity to work in a fast-paced, interdisciplinary team environment.
  • Be part of a new role specifically designed for owning evaluation processes.
  • Contribute directly to innovative healthcare solutions with tangible impacts.
  • Flexibility to define your methods and practices in an ambiguous environment.
Full Job Description
Your team

We are a small, interdisciplinary team - engineers, data scientists, designers, product managers, and clinicians - building Nuna's AI health coach. Our products are only as good as the care and science behind them, and your piece is how we know the coach is safe and working. You'll own the evaluation system for the team building the coach: a data scientist partners with you on the science, clinicians and designers supply the ground truth, and the engineers shipping the agents depend on the signal you produce to decide what ships.

The role

This is a net-new, build-first role for someone who wants to own how we evaluate our AI agents end to end. You'll build the harnesses, datasets, judges, and release gates that tell us whether the coach is safe and good, and you'll own both that infrastructure and the evals that run on it. This is not a test-execution role - you write the code and own the system, rather than running tests someone else designed. You'll make the day-to-day calls on standards, methods, and trade-offs, often with incomplete information and the freedom to define the right answer yourself. Comfort with ambiguity is part of the job.

What you'll do
  • Build testing harnesses and evaluation infrastructure for our agentic products and our internal agentic tooling
  • Own our evals end to end - both the architecture and the content - with support from data science and clinical partners
  • Make every agentic deployment run through the testing apparatus before it ships, and own the release gates that keep unsafe or low-quality behavior from reaching patients
  • Build the ground truth, judges, and metrics, and validate that the evaluation itself can be trusted: calibration to human labels, reliability, and honest confidence on every number, in partnership with our data scientist
  • Build functional tooling for labeling and review workflows, so clinicians, coaches, and designers can author and review evaluation scenarios without an engineer in the loop
  • Help close the loop from evaluation results to model and prompt refinement, working toward systems that iterate safely with less human hand-holding
What we're looking for
  • Significant experience building and shipping reliable production systems and tooling
  • Deep understanding of how to evaluate AI systems - LLM-as-judge, red-teaming and adversarial testing, synthetic scenario generation, and multi-turn and agentic evaluation - and a clear sense of how evals themselves fail. You've deployed evals and automated AI tooling in production, not just prototyped them
  • A testing mindset applied to building the measurement system, not running tests against a spec: adversarial instinct, coverage thinking, regression discipline, and documentation others can build on
  • You use AI in your daily work and build tools that make the people around you more effective
  • Enough fluency in statistics and experimental design to partner with a data scientist on calibration and reliability
  • Can design the workflow and build a functional UI for non-engineers like clinicians and labelers
  • A genuine interest in improving healthcare alongside an interdisciplinary team, with the judgment to tell a launch-blocking issue from a nice-to-have

Bonus Points
  • Experience in healthcare or another regulated, high-trust domain, and familiarity with the regulatory landscape
  • Hands-on experience with the eval tooling ecosystem (LangSmith, Braintrust, DeepEval, Ragas, Promptfoo, or similar)
  • Red-teaming or AI safety experience - prompt injection, jailbreaks, adversarial and stress testing
  • Experience with automated, eval-driven model or prompt optimization
  • You've built in an early-stage or fast-moving environment

#LI-LM1

About Nuna Incorporated

Nuna is a healthcare technology company that provides data and analytics solutions to payers, providers, and government agencies. The company's platform aggregates and analyzes healthcare data to help organizations improve patient outcomes and reduce costs. Nuna's customers include Medicaid agencies, health plans, and hospital systems. The company was founded in 2010 and is headquartered in San Francisco, California.
Learn more about Nuna Incorporated
Size
200 employees
Industry
Founded
2010

Similar Jobs

More Jobs at Nuna Incorporated

  • Nuna Incorporated
    Lead Software Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person
  • Nuna Incorporated
    Software Engineer
    $120K — $145K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person
  • Nuna Incorporated
    Lead Data Engineer
    $130K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person
  • Nuna Incorporated
    Product Designer
    $100K — $140K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person
  • Nuna Incorporated
    Senior Software Engineer
    $130K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Healthcare
    In-Person

More Healthcare Jobs

Find similar Software Engineer, AI Evaluation jobs: