Job Summary
We are seeking a Senior AI Engineer with strong hands-on experience building and deploying production-grade agentic AI systems. The ideal candidate will have deep expertise in Python, LangChain/LangGraph, GenAI prompt engineering, PyTorch, OpenAI technologies, and modern testing and DevOps practices. This role focuses on designing reliable AI agents capable of serving real users at enterprise scale, with emphasis on production delivery, measurable outcomes, scalability, latency, reliability, safety, and observability.
Key Responsibilities
• Design and build production agentic workflows for customer-facing use cases, including multi-turn conversations, tool calling, retrieval, escalation, and human handoff.
• Own agent orchestration using LangGraph, LangChain, or equivalent frameworks, including state management, guardrails, error handling, and failure recovery.
• Integrate AI agents with Java/Spring microservices, APIs, event streams, and enterprise service layers.
• Build and optimize evaluation frameworks, including automated evaluations, regression suites, trace analysis, and A/B testing to measure agent quality.
• Develop inference and serving capabilities, including model routing, prompt and context management, and self-hosted model serving using vLLM or comparable technologies.
• Integrate commercial LLM APIs such as OpenAI, Anthropic, Amazon Bedrock, or Google Vertex AI.
• Design and implement Retrieval-Augmented Generation (RAG) pipelines and tool/function-calling workflows.
• Apply prompt engineering and context engineering techniques to improve agent accuracy, reliability, and user experience.
• Implement robust testing strategies for AI workflows, APIs, integrations, and production systems.
• Troubleshoot model behavior, agent failures, distributed-system issues, and production performance problems.
• Implement security and compliance controls for regulated environments, including PII handling, auditability, deterministic fallbacks, and observability.
• Build and maintain CI/CD pipelines and DevOps processes supporting AI applications and services.
• Monitor production AI systems and optimize latency, reliability, scalability, and operational performance.
• Collaborate with engineering, product, security, compliance, and architecture teams to deliver production-ready AI capabilities.
Required Qualifications
• 7+ years of software engineering experience with senior or lead-level ownership of production systems.
• Proven hands-on experience delivering LLM-based agentic systems into production.
• Strong Python development experience for AI, ML, automation, and agent applications.
• Strong Java development experience for integration with enterprise services and microservices.
• Hands-on expertise with LangChain and LangGraph or equivalent agent frameworks.
• Strong GenAI prompt engineering and context engineering experience.
• Experience building RAG pipelines and implementing tool/function calling.
• Working knowledge of PyTorch and machine learning model fundamentals, including model behavior, debugging, and fine-tuning concepts.
• Experience with LLM inference and serving technologies such as vLLM, TGI, or TensorRT-LLM.
• Experience integrating commercial LLM APIs, including OpenAI, Anthropic, Bedrock, or Vertex AI.
• Strong understanding of distributed systems, APIs, queues, caching, observability, and production reliability.
• Experience with automated testing, regression testing, CI/CD, DevOps, and production deployment practices.
• Strong analytical and troubleshooting skills with the ability to diagnose complex AI and software-system failures.
• Experience designing secure and auditable AI solutions for enterprise environments.
Preferred Qualifications
• Experience with agent-builder platforms such as Sierra or Decagon.
• Experience building voice agents, real-time AI systems, streaming inference, or contact-center integrations.
• Experience in regulated industries such as financial services, healthcare, or insurance.
• Experience with model risk management, compliance reviews, audit trails, and AI governance.
• Experience with evaluation frameworks such as LangSmith, Braintrust, or custom evaluation harnesses.
• Experience with LLM safety, guardrails, and responsible AI tooling.
• Experience with cloud platforms such as AWS, Azure, or Google Cloud.
• Experience with Kubernetes, Docker, infrastructure automation, and cloud-native architectures.
• Experience mentoring engineers and leading technical design and architecture discussions.
Candidate Profile
• Strong candidates should lead with delivery details, clearly demonstrating what they built, the scale at which it operated, and whether the solution reached production.
• Resumes should demonstrate at least one shipped AI agent or retrieval-based system, including the framework used, real-user adoption, and measurable business or technical outcomes.
• Candidates should provide concrete production-scale indicators such as transaction volumes, latency, uptime, throughput, or user counts.
• Candidates should demonstrate the ability to distinguish production-grade AI systems from proof-of-concept or experimental implementations.
• Strong communication, ownership, curiosity, and ability to operate in a fast-paced engineering environment are essential.
Certifications
• Relevant cloud, AI/ML, Kubernetes, or DevOps certifications are a plus but not required.