WHAT YOU'LL DOIn this role, you will build the intelligence behind the product that gets stuff built. Permitting runs on messy inputs-scanned plan sets, jurisdiction code, reviewer comments, application forms that differ in every city-and turning that into something fast, structured, and trustworthy is the core technical problem at Pulley. As a senior-level AI engineer, you will:
- Own AI-powered features end-to-end-from talking to users and defining what "correct" means for a permitting workflow, through prompt and pipeline design, evals, deployment, and iteration in production
- Turn unstructured permitting documents, city regulations, and jurisdiction workflows into structured, reliable outputs-extraction, classification, retrieval, and agentic workflows over documents that were never designed to be machine-readable
- Build the evaluation and observability foundation that lets us ship LLM-powered features with confidence: define ground truth, measure quality and regressions, and know when a model change is actually an improvement
- Build with AI agents as a daily practice-directing, reviewing, and shipping agent-driven work at high velocity while owning the quality bar
- Make technical and product decisions that have direct impact on our customers and their projects
- Raise the bar for the engineers around you in how they build with LLMs, through design review, mentorship, and the standards you set in your own work
WHO YOU ARE- You thrive in ambiguity-you'd rather define the right problem than execute a spec, and you're energized rather than paralyzed when the path isn't laid out
- You're product-minded: you care whether the thing you built actually solved the customer's problem, and you'll talk to users to find out
- You're rigorous about what "working" means-you don't trust a demo, you trust an eval, and you build the measurement before you build the feature
- You have strong opinions about quality and velocity and don't treat them as a tradeoff-you look for the tools, abstractions, and processes that buy both
- You default to ownership: when something is broken or missing, your instinct is to fix it, not to file it as someone else's problem
NEED TO HAVE- 4+ years of software engineering experience, with a substantial portion building production LLM or ML systems
- Track record of owning an LLM-powered product surface end-to-end: requirements through production, including the unglamorous parts-data quality, eval design, cost and latency, failure handling
- Deep hands-on experience with large language models in production-prompting, retrieval-augmented generation, structured extraction, tool use and agentic workflows, and knowing when each is the wrong tool
- Experience designing evals and otherwise making LLM-powered features reliable in production
- Real experience building with AI coding agents-not just autocomplete; you've shipped work where agents did substantial implementation under your direction
- Ability to architect durable systems while making pragmatic tradeoffs
- Based in the San Francisco Bay Area and willing to work in person 4 days a week
NICE TO HAVES- Experience with document understanding at scale-OCR, layout-aware parsing, or vision-language models over scanned PDFs, drawings, or forms
- Experience fine-tuning models or building data pipelines to produce training and eval sets from real-world usage
- Experience in construction tech, govtech, proptech, or another domain where the hard part is messy real-world documents and processes
- Experience with modern full-stack development-we use TypeScript, React, and Google Cloud-and an appetite for working in the application code that puts AI features in front of users
- Startup experience at the stage where you helped build the team, not just the product
- Experience mentoring engineers or leading technical direction across teams