5+ years building production ML or backend systems, focusing on LLM or agent applications from prototype to production.
Strong understanding of agent design principles, including planning, reasoning, tool use, orchestration, and memory.
Experience in developing rigorous evaluation systems as well as execution tracing and observability for agents.
Familiarity with OpenAI Responses API and understanding of server- versus client-side execution.
Responsibilities
Design agent architectures for planning, reasoning, tool use, memory, and integration with external systems and data.
Improve reliability for long-running, multi-step tasks, including failure recovery.
Build evolutionary loops for agents to enhance their performance over time.
Develop robust evaluations to measure agent performance without manipulation.
Make practical trade-offs among quality, latency, cost, and complexity.
Benefits
Flexible work options including in-person collaboration in the Bay Area and global remote team dynamics.
Unique Adaption Passport offering annual travel stipend to visit a new country to foster personal and professional growth.
Weekly lunch stipend for take-out meals or grocery delivery.
Comprehensive medical benefits paired with generous paid time off for employee well-being.
Full Job Description
The Role
You'll build the agent systems at the core of our product. These systems turn customer goals into reliable, multi-step execution across real tools and services.
This is not about building demos. You'll work on agents that operate under real constraints: incomplete information, external failures, limited budgets, and unpredictable traffic. You'll own how they plan, use tools, recover from errors, and improve over time.
Responsibilities
Design agent architectures for planning, reasoning, tool use, memory, and integration with external systems and data.
Improve reliability on long-running, multi-step tasks, including failure recovery.
Build the loops that let agents improve with real use, so performance compounds instead of staying frozen.
Develop evaluations that measure agent performance and resist being gamed.
Make practical tradeoffs between quality, latency, cost, and complexity.
Qualifications
5+ years building production ML or backend systems, including taking LLM or agent applications from prototype to production.
Strong understanding of agent design: planning, reasoning, tool use, orchestration, and memory.
Experience building rigorous evaluation systems, plus execution tracing and observability for agents, with a focus on reproducibility.
Familiarity with the OpenAI Responses API, MCP, and server- versus client-side execution.
Above all, we're looking for great teammates who make work feel lighter and aren't afraid to go out on a limb with bold ideas. You don't need to be perfect, but you do need to be adaptable. We encourage you to apply, even if you don't check every box.
Benefits
Flexible work: In-person collaboration in the Bay Area, a distributed global-first team, and team offsites.
Adaption Passport: Annual travel stipend to explore a country you've never visited. We're building intelligence that evolves alongside you, so we encourage you to keep expanding your horizons.
Lunch Stipend: Weekly meal allowance for take-out or grocery delivery.
Well-Being: Comprehensive medical benefits and generous paid time off.