Role Overview:This role involves building and scaling a production multi-agent AI platform designed to serve thousands of internal users across various business units. The platform will operate with a monthly release cadence, addressing real-world challenges related to user experience, latency, and cost efficiency.
Key Responsibilities: - Develop an LLM-driven orchestrator for routing user intent across a portfolio of specialized agents, managing delegation, memory, response validation, and capability discovery.
- Implement an agent selection layer utilizing hybrid retrieval (vector RAG over a capability registry) and closed-set LLM selection with JSON-schema-constrained outputs.
- Build a multi-agent SDK/gateway using FastAPI, hosting multiple agents behind path-prefix routing, managing per-agent tool registries, and session-scoped conversational context.
- Create tool-driven agents with 15-30 dynamically composed tools by an LLM, ensuring tool contracts, guardrails, and evaluation.
- Design and implement a data API layer with parameterized endpoints between agents and databases, ensuring LLMs do not directly interact with databases.
- Facilitate partner-team onboarding, including versioned A2A contracts, bring-your-own-agent registration, and auto re-embedding.
Required Skills: - Proficiency in production LLM systems, including RAG, tool/function-calling loops, structured outputs, hallucination guards, and closed-set selection.
- Experience with multi-agent orchestration, A2A protocols, session affinity, human-in-the-loop gating, kill switches, and graceful degradation.
- Expertise in vector search and embeddings at scale, achieving sub-second retrieval over thousands of documents.
- Knowledge of evaluation & safety practices: PII/PHI masking, audit trails, feedback-loop instrumentation, and offline + online evaluation.
- Strong programming skills in Python 3.11+, FastAPI, async I/O, and Pydantic.
- Familiarity with modern LLM stacks such as Gemini, GPT, Claude, and agent frameworks like LangGraph and other Agent SDKs.
- Cloud platform experience (GCP or AWS), including Kubernetes, object storage, workflow orchestration, and Vertex/Bedrock-class services.
- Database and caching experience with Redis, MongoDB, Oracle/Postgres, and knowledge of SSO + RBAC.
- Understanding of observability tools and practices: Prometheus, structured JSON logs, per-decision audit trails, and p95 latency SLOs.
- Skills in Machine Learning, Artificial Intelligence (AI), and Generative AI.
Qualifications: - 4-6 years of experience in a relevant engineering role.