Job Summary
We are seeking a Senior AI Engineer to design and build production-grade agentic systems for customer-facing use cases. This role will focus on developing reliable AI agents capable of multi-turn conversation, tool calling, retrieval, escalation, and human handoff within a regulated, high-trust environment. The ideal candidate will have proven experience shipping LLM-based agentic systems to production and strong expertise across Python, Java, LangChain, LangGraph, RAG, model serving, evaluation, and distributed systems.
Key Responsibilities
• Design and build production agentic workflows for customer-facing use cases, including multi-turn conversation, tool calling, retrieval, escalation, and human handoff.
• Own agent orchestration end to end using LangGraph, LangChain, or equivalent frameworks, including state management, guardrails, and failure recovery.
• Integrate agents with core service layers, including Java/Spring microservices, APIs, and event streams.
• Build and tune evaluation loops, including automated evaluations, regression suites, trace analysis, and A/B measurement of agent quality.
• Work on inference and serving layers, including model routing, prompt and context management, and self-hosted serving with vLLM alongside commercial LLM APIs.
• Implement solutions for regulated, high-trust environments, including PII handling, auditability, deterministic fallbacks, and observability.
Required Qualifications
• 7+ years of software engineering experience, with senior or lead-level ownership of production systems.
• Proven hands-on experience delivering LLM-based agentic systems in production, including the ability to explain architecture, failure modes, and evaluation strategies for shipped systems.
• Strong Python experience as a primary agent and ML development language and Java experience for service integration.
• Advanced experience with LangChain, LangGraph, prompt and context engineering, tool/function calling, and RAG pipelines.
• Working knowledge of PyTorch and model fundamentals, including the ability to fine-tune models, debug model behavior, and evaluate technical tradeoffs.
• Experience with LLM serving and inference using vLLM or comparable technologies such as TGI or TensorRT-LLM.
• Experience working with commercial LLM APIs such as OpenAI, Anthropic, Bedrock, or Vertex.
• Strong distributed systems fundamentals, including APIs, queues, caching, observability, CI/CD, and production engineering practices.
• Advanced experience with GenAI prompt engineering.
• Advanced experience with OpenAI engineering.
• Advanced experience with testing, DevOps, and CI/CD practices.
Preferred Qualifications
• Hands-on experience with agent-builder platforms such as Sierra or Decagon, including deploying, extending, or evaluating platforms for customer-service use cases.
• Experience with voice agents, real-time or streaming inference, or contact-center integrations.
• Experience working in regulated industries such as finance or healthcare, including model risk, compliance review, and audit trails.
• Experience with evaluation frameworks such as LangSmith, Braintrust, or custom evaluation harnesses.
• Experience with LLM safety and guardrail tooling.