About the RoleThe Foundation Agents team stewards the long-horizon agents powering the Fieldguide AI platform. We work at the frontier of AI product development: agent knowledge, evaluations, and improving quality and reliability at scale. As a Senior Software Engineer, Agents, you'll take ownership of how the team measures and improves agent quality, and help drive the platform forward.
What You'll Own- Evals strategy and error-analysis practice, shaping how the team measures and improves agent quality
- Design and build agent knowledge and evaluation infrastructure for Fieldguide's long-horizon agents
- Lead error analysis on agent behavior, turning findings into concrete platform-level reliability improvements
- Build and harden backend systems that support agent execution, evaluation, and monitoring at scale
- Drive the AI platform's reliability and quality roadmap forward, working closely with the broader AI team
- Mentor engineers on the team, raising the bar on eval rigor and error-analysis practice
Who You Are- You've built AI products end-to-end, with real ownership over agent quality outcomes
- You think in evals and error analysis as a discipline; you dig into why an agent failed and fix the systemic cause
- You're strong in the backend and comfortable owning platform-level systems
- You're motivated by long-horizon agents that do real work in production
- You multiply the people around you, not just your own output
ExperienceMust-have:- 1+ years working specifically on agents
- Strong experience working on an AI platform
- Demonstrated experience building evals and performing error analysis
- Backend engineering experience
Nice-to-have:- Strong platform engineering skills
- Frontend experience
- Distributed systems experience
What Should Excite You- Long-horizon agents: Working on agents that do real, sustained work
- Evaluation as a craft: Evals and error analysis are core to how this team improves quality, not an afterthought
- Platform-level impact: Your work shapes the reliability and quality of every agent built on top of it
- Frontier problems: You're working on open problems in agent reliability that don't have established playbooks yet
Benefits- Competitive compensation with equity
- Comprehensive health and wellness benefits
- Flexible time off and work schedules
- Technology reimbursements
- 401(k) plan
- Twice-yearly in-person offsites across the U.S.
- Wellness benefits starting on your first day