AI Engineer, Evals & Agent Quality

Town, Inc.

$150K — $180K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in AI model evaluation or related field
  • Strong background in offline/online quality measurement systems
  • Hands-on experience with existing evaluation frameworks and tools
  • Ability to analyze and reason about model routing decisions
  • Proven track record of deploying fixes and enhancements in a technical environment
  • Senior or staff-level engineer with greenfield project experience

Responsibilities

  • Build a comprehensive evaluation framework for measuring assistant quality across multiple touchpoints
  • Establish and maintain golden datasets alongside a continual validation loop
  • Develop model routing strategies and online evaluation tools to optimize model performance
  • Ensure every system change is quantifiable to facilitate rapid iteration
  • Collaborate with engineering teams to instrument quality measurements and implement improvements

Benefits

  • Five days a week in-person work environment in a vibrant San Francisco office
  • Opportunity to work on groundbreaking AI technology
  • Autonomy and ownership over foundational systems and frameworks
  • Collaborative culture partnering with skilled engineers
  • Focus on the measurable impact of work with immediate feedback on system performance
Full Job Description
About the role

Town is building the most personalized, most capable AI assistant for everyone - one that knows you deeply, works across every tool you use, and gets sharper over time. Building the best assistant means proving it's the best: every model, prompt, and system change has to be measurably better, on every surface it touches.

That's what you'll own. You'll build the evals and quality systems that turn assistant performance into numbers the whole team can trust, measuring and improving the full multi-step trajectory the assistant takes to do real work. You'll build the model routing that puts the right model in the right place balancing cost, quality, and speed.

This is a foundational, 01 build with ownership to match: the eval framework, the golden datasets and labeling loop, model routing, and online measurement, and you set the bar for what "best" means at Town.

What you'll do
  • Build a generalized eval system that measures assistant quality across every surface it touches - and, crucially, across multi-step agent trajectories.
  • Stand up golden datasets and the labeling loop that keeps them up to date and constantly checking to validate improvements and avoid regressions.
  • Build model routing and online evaluation tooling to help us learn and route to the best models.
  • Make every prompt and system change measurable, so the team can move fast without breaking what works.
  • Partner with engineers across the product to instrument quality and close the loop from signal to fix.
You might thrive here if you...
  • Have built or owned LLM eval systems, or offline/online quality measurement at scale.
  • Think rigorously about measurement. Maybe that came from an MLE or applied-ML background, maybe not, the instinct for how to measure "better" matters more than the exact pedigree.
  • Know the eval landscape hands-on, off-the-shelf tooling and eval frameworks, and have opinions on what to reach for when.
  • Are comfortable reasoning about model routing and the tradeoffs between models.
  • Ship the fixes, not just the dashboards and metrics.
  • Are a senior or staff engineer comfortable in greenfield, where the system doesn't exist yet.


Location

San Francisco, CA. Five days a week in person at our Financial District office.

Similar Jobs

More Jobs at Town, Inc.

More Consumer Technology Jobs

Find similar AI Engineer, Evals & Agent Quality jobs: