Full Job Description
We9re looking for an AI Engineer who can build production-grade AI systems end-to-end - from prototype to pipeline to product - with the ownership and urgency of a startup culture.
This is not a "wire up a prompt chain and move on" role. You9ll own core pieces of the AI stack that power Hilbert9s demand intelligence platform - designing agent architectures, building evaluation systems, and making hard tradeoffs between accuracy, latency, and cost in production. You9ll ship fast in conditions where the spec is evolving, and communicate what you9re building (and why) with clarity to the rest of the team. If you think in systems, have opinions about how agentic workflows should actually work, and want to build AI products that drive real enterprise outcomes, we want to meet you.
What you9ll own first:
Evaluation and testing for our agents. Hilbert9s agents are in production with enterprise customers today. Before we expand what they do, we need to know, reproducibly, when a change makes them better or worse. You9ll own designing the eval harness, defining what "correct" means for a multi-step agent trajectory, building the regression gates that run before anything ships, and turning production failures into test cases. From there, the scope widens into retrieval, orchestration, and execution across the AI stack.
What you9ll do:
3 Own the evaluation layer for our agents - harnesses, metrics, golden datasets, regression gates, and human-in-the-loop review.
3 Architect and implement agent workflows using LangChain, LangGraph, or equivalent; state memory, routing, tools registries and recovery paths.
3 Own systems from experimentation through production and operate in production: tracing, monitoring, latency, cost-per-task budgeting, on-call for what you build.
3 Diagnose real production failures: hallucination, tool misuse, retrieval misses, silent degradation and turn each into a durable fix and a test.
3 Set the technical standard for agent work at Hilbert: review designs, define patterns others build on, and raise the bar across the team.
3 Collaborate closely with the founding team and cross-functional partners - communicating tradeoffs, progress, and technical decisions with clarity.
3 Make pragmatic engineering decisions under ambiguity-ship, learn, iterate.
Our Current Hurdles
These are the kinds of problems you9ll walk into on day one:
3 Intelligent retrieval across heterogeneous approaches - our agents need the right information at exactly the right moment. The challenge isn9t picking one retrieval method; it9s combining RAG, graph-based retrieval, and other approaches into a unified strategy that fetches the most relevant content precisely when the agent needs it - no more, no less.
3 Agentic workflows that solve real-world problems - it9s building workflows robust enough to handle the unexpected. When an agent hits an edge case, missing data, or a situation it wasn9t explicitly designed for, it needs to reason through it - leveraging available context, escalating to a human when it can9t, and never silently failing.
3 Evaluation beyond vibes - we need systematic, reproducible evals that actually predict real-world performance. If you9ve built custom evaluators for RAG or agent workflows, we want to talk.
3 Execution and real-world integration - an agent that only surfaces insights isn9t enough. We9re building systems where agents take action - integrating with external platforms, executing workflows, and doing real work with the information they have, combined with human-in-the-loop checkpoints that keep enterprise trust intact.
WHO THRIVES IN THIS ROLE
We care about what you9ve shipped and how you think, but as a Senior, we expect you to demonstrate 6+ years of building production software, with at least 2 of those on LLM or agentic systems that real users depend on.
Must-haves
3 6+ years of production software engineering: APIs, services, data infrastructure. You9ve owned code that other people depended on, with tests, CI/CD, and on-call attached.
3 2+ years building LLM or agent systems that shipped to productions watching real users adopt. Not internal demos, not prototypes. Be ready to talk about what broke in production and what you did about it.
3 Hands on with LangChain, LangGraph, or equivalent agent/orchestration frameworks. You9ve built with them, hit their limits, and worked around them. You operate beyond just following tutorials.
3 You communicate with clarity and conviction. You can explain a technical decision to a non-technical founder and debate architecture tradeoffs with a senior engineer. Communication is not a nice-to-have here - it9s core to the role.
3 You take ownership. You don9t wait for tickets. You see what needs to be built, raise your hand, and ship it.
3 You thrive in ambiguity. AI products evolve fast. Requirements change. You9re energized by figuring it out.
3 You move at startup speed. Without waiting for permission, and you know which decisions deserve a day of thought and which deserve an hour.
Strong pluses:
3 Retrieval-augmented generation (RAG) at depth: Hybrid and graph retrieval, chunking and embedding strategy, ranking, grounding
3 Observability for LLM systems (Langfuse, OpenTelemetry, or equivalent) and cost/latency optimization
3 MCP, tool-calling frameworks, structured output and constrained decoding
3 Experience at early-stage startups or high-growth environments where you wore multiple hats
You might be:
A backend engineer who went deep on LLMs and never looked back. An ML engineer who realized they love building products, not just models. A startup CTO who wants to go deep on AI at a company where the stack is the product. Someone who9s been hacking on agents and pipelines nights and weekends and wants to do it full-time with real enterprise stakes. What matters: you ship, you own it, and you communicate like a teammate - not a silo.
Location
San Francisco, with occasional travel for team meets, offsites or customer engagements.
The Hiring Journey
Short form 12 Intro call 12 Technical working session 12 Team conversations 12 Offer
Why join us
At Hilbert, we move fast, work collaboratively, and give people real ownership over their impact. As a fast-growing company, there9s no shortage of room to grow - you9ll take on new challenges quickly and shape the path as you go.
On top of that, here9s how we take care of our team:
Health & wellness. We cover 100% of your Health, Dental, and Vision premiums - with the option to add coverage for dependents or a spouse at additional cost.
Financial future. Plan ahead with our 401k through Human Interest, including an employer match that vests immediately.
Time to recharge. Enjoy generous PTO so you can rest, travel, and spend time on what matters most outside of work.
Commuter support. Based in SF? We help cover your commuting costs.
Lunch, covered. In the SF office, we provide lunch vouchers via DoorDash.
And many more. We9re always looking for new ways to support our team. Come build the future with us, and grow alongside a company that9s investing in yours.