About the RoleYou'll join a genuine 01 team on the ground floor of one of the company's biggest new bets. This seat is specifically
product-focused: you'll own agent quality, ship agents that do real audit work, and work alongside practitioners.
Depending on your experience and what you're looking to own, you may join building within a major agent area, owning one end-to-end, or setting technical direction for agentic audit work across the team. We're hiring across all levels and will calibrate during interviews based on scope and demonstrated experience.
What You'll Do- Make agent judgment repeatable: run error analysis on real testing data and turn findings into concrete fixes
- Tradeoffs such as quality/latency/cost across a long multi-phase run
- Build structured-output pipelines that turn model output into real audit artifacts
- Take ambiguous problem statements and turn them into a plan, a shipped feature, and a clear read on what was cut and why
- Work directly with an embedded subject matter expert and with design-partner firms, turning their feedback into agent changes within days
- Expand agent coverage into new controls and new areas of internal audit
Who You Are (All Levels)- Product-minded and full-stack: you've shipped LLM-backed features to production against real users, and you measure yourself on whether they got used
- You're fluent in evals and error analysis, and you apply them in service of shipping something practitioners trust
- You have real opinions on model selection, prompting, and orchestration tradeoffs, and you can defend them with evidence rather than vibes
- Energized by 01 work: you'd rather define the problem than inherit a spec, and you don't stall on ambiguity
- Strong instincts for human-in-the-loop design
- A genuine team player across the organization, not just within engineering: you'll work daily with PM, design, domain experts, and customer-facing teams, and you treat that as the best part of the job
- Ship fast without leaving a mess: your code is reviewable, tested where it counts, and instrumented
- Able to internalize a hard domain fast. You don't need to know SOX today, but you'll understand it well enough to make the right product calls
Higher-Level ResponsibilitiesAt the Senior level, you may:
- Own a major agent area end-to-end, from how the agent reasons about a class of controls through to the artifact a reviewer signs
- Set the evals and error-analysis practice for the team's agent work, and decide what evidence justifies shipping a change or rolling it back
- Collaborate with PMs and designers to shape roadmaps and define architectural tradeoffs, including where the agent acts and where the auditor decides
- Own the harder model and orchestration judgment calls across a long multi-phase run
- Mentor other engineers and raise the bar on 01 execution and applied eval rigor
At the Staff level, you may:
- Drive agent initiatives that reach beyond Internal Audit and influence how agents are built across Fieldguide
- Set and champion engineering standards for agent reliability, reproducibility, and defensibility
- Partner with engineering and product leadership to define long-term technical strategy for agentic audit work
- Serve as a trusted advisor to leaders across Engineering, Product, and Design
- Represent Fieldguide externally through writing, speaking, and open-source contributions
ExperienceMust-have:- Shipped LLM-backed product features to production against real users
- Applied AI skillset: evals, error analysis, and model-selection decisions you owned and can explain
- Comfortable full-stack, with enough backend depth to work in agent orchestration
- Autonomy working from an ambiguous spec
- A collaborative mode that works across PM, design, and domain experts
Nice-to-have:- Python, TypeScript, React, Postgres, Hasura, GraphQL
- Temporal or comparable durable-execution / workflow orchestration
- Hands-on eval experience (Langfuse, Braintrust, LangSmith, Arize Phoenix, or comparable)
- Structured-output work including schema contracts, generating real artifacts from model output
- Startup experience, as a founder or as an early engineer
- A 01 track record: things you started where no scaffolding existed
- Experience working directly with customers, and comfort being in the room when they use what you built
- Background in internal audit, SOX, accounting, or another regulated domain
- Document processing, including PDF and Excel manipulation and annotation
Not a fit if:- Prompt engineering is your whole skill set
- Your agent work never carried production traffic
- You want to own eval methodology or the evaluation harness itself rather than ship product features (better fit on Foundation Agents)
- You need a fully specified ticket to start
- You'd rather not be in the room with customers and domain experts
What Should Excite You- 01 on the biggest bet: You're building the agent and the product from scratch, on the ground floor of where the company is going
- Repeatable judgment: Making an agent reach the same defensible conclusion twice, in a domain where ground truth requires expert judgment
- Real audit stakes: Your work directly affects what firms put in front of their clients, and what a reviewer is willing to sign
- Customer proximity: Design-partner firms and an embedded SOX expert use what you ship within days of it landing
- Human-in-the-loop design: Deciding where the agent acts and where the auditor decides, on work that genuinely matters
- High trust, high autonomy: You're given ambiguous problems and trusted to define the plan
Benefits- Competitive compensation with equity
- Comprehensive health and wellness benefits
- Flexible time off and work schedules
- Technology reimbursements
- 401(k) plan
- Twice-yearly in-person offsites across the U.S.
- Wellness benefits starting on your first day