About the RoleThis is a measurement-first role that owns the quality and efficiency of Cortex Code end to end: how good the agent is, how much it costs to run, and how reliably it behaves in production. You will take agents from research capability to real, measurable user value - turning fuzzy "the agent feels worse" signals into hard metrics, running the experiments that move them, and shipping the changes that stick. You will work on a small, high-powered modeling and infrastructure team where your work reaches every developer building on Snowflake.
What you will do in this role- Take agents from prototype to production: design and refine agent behaviors for real coding and data-engineering workflows, and make them reliable enough to depend on.
- Own agent quality end to end: build eval harnesses, diagnose failure modes from real agent trajectories, and run experiments that hillclimb the metrics that matter.
- Drive down cost and latency without regressing quality: prompt caching, context compaction, tool-result offloading, and cheaper model routing for sub-tasks. At Snowflake scale, efficiency is user value.
- Debug production failures and systematically increase robustness: close the loop from a customer's broken run back to a fix and a regression test.
- Build the data pipelines that feed real-world task insights back into evals and modeling.
- Onboard and bake off frontier models: measure their strengths and failure modes, and decide where each belongs in the product.
- Partner with product and infra: shape user-facing agent behavior, and set the metrics and standards that define a successfully completed complex task.
Requirements- Bachelor's degree in Computer Science, Engineering, Statistics, or a related field. Master's or higher preferred but not a requirement.
- 6+ years of experience shipping software in production, including AI/LLM features.
- Fluency in at least one of Python, TypeScript, or Go, and willingness to work across all three
- A measurement-first, systems-thinking instinct: you optimize for user outcomes over isolated metrics, and you can design an eval that is not fooling you (sampling, ground-truth quality, leakage, noise).
- Comfort debugging complex, unpredictable, real-world failures.
- Strong communication skills: you can make a quality or cost result legible to engineers, product, and leadership, and collaborate effectively in a team environment.
Nice to have- Deep hands-on experience with agentic coding tools and real intuition for model strengths, failure modes, and prompting limits.
- Prior work on eval harnesses, LLM observability, or safety/guardrails in production.
- Background in data engineering, data modeling, analytics, retrieval/RAG, or semantic layers, which is highly relevant for data-centric coding agents.
- Experience working with large-scale datasets or production system logs.
You may be a particularly good fit if you- Are a power user of modern coding agents and want to turn that intuition into systematic measurement and improvement.
- Have built and owned complex systems - pipelines, orchestration, or software with substantial state, branching logic, and operational requirements.
- Thrive in high-intensity environments with short feedback loops and high standards for rigor.
- Take problems to completion independently: you don't stop at a prototype; you care about production reliability and clear metrics.
- Are genuinely bothered by numbers that do not reconcile.
Snowflake is growing fast, and we're scaling our team to help enable and accelerate our growth. We are looking for people who share our values, challenge ordinary thinking, and push the pace of innovation while building a future for themselves and Snowflake.
How do you want to make your impact?
For jobs located in the United States, please visit the job posting on the Snowflake Careers Site for salary and benefits information: careers.snowflake.com