Job Summary
We are seeking an Artificial Intelligence Engineer with strong experience building production-grade Python services and multi-agent LLM systems. The ideal candidate will have expert-level knowledge of asynchronous Python, FastAPI/ASGI, LLMOps, agent evaluation, and Google Cloud/Vertex AI. This role will focus on building scalable AI solutions, implementing agent orchestration and guardrails, and supporting production-grade observability, evaluation, and deployment practices.
Key Responsibilities
• Build and maintain production-grade Python services using asynchronous Python, FastAPI, and ASGI.
• Design, develop, and deploy multi-agent LLM systems using technologies such as Google ADK, LangGraph, A2A, or comparable frameworks.
• Implement agent tool use, orchestration, workflows, guardrails, and production-ready AI capabilities.
• Develop and maintain LLMOps and agent evaluation processes using tools such as Phoenix/Arize, Vertex AI evaluation, and OpenTelemetry tracing.
• Define and monitor AI quality, safety, performance, and evaluation metrics.
• Design and implement AI solutions using Google Cloud Platform and Vertex AI, including search, retrieval, and model serving.
• Develop prompt engineering strategies and structured-output schemas using Pydantic.
• Build and support containerized services and production deployment environments.
• Develop and maintain CI/CD pipelines using Harness or similar platforms.
• Implement observability solutions for AI and cloud-native services at scale.
• Design and implement RAG and vector search solutions as needed.
• Support scalable MySQL database implementations and enterprise authentication using OAuth2 and scopes.
• Contribute to MLOps experiment tracking and AI application lifecycle management.
• Collaborate with engineering and architecture teams to define scalable and maintainable AI solutions.
• Mentor engineers and contribute to technical architecture and engineering best practices.
• Support modernization initiatives involving reactive technologies, including Java, Spring WebFlux, Project Reactor, and reactive programming patterns.
Required Qualifications
• Production experience building Python services.
• Expert-level experience with asynchronous Python and FastAPI/ASGI.
• Hands-on experience building multi-agent LLM systems using Google ADK, LangGraph, A2A, or comparable technologies.
• Experience implementing LLM agent tool use, orchestration, and guardrails.
• Proven experience with LLMOps and agent evaluation.
• Experience with Phoenix/Arize, Vertex AI evaluation, OpenTelemetry tracing, or comparable evaluation and observability technologies.
• Strong experience with Google Cloud Platform and Vertex AI, including search, retrieval, and model serving.
• Production experience with containerized services.
• Experience with CI/CD platforms such as Harness or similar tools.
• Strong understanding of observability practices for production-scale services.
• Experience with prompt engineering and structured-output/schema design using Pydantic.
• Strong understanding of AI/ML application development and production deployment.
Preferred Qualifications
• Experience with RAG and vector search technologies.
• Experience working with MySQL at scale.
• Experience with enterprise authentication, including OAuth2 and scopes.
• Experience with MLOps experiment tracking.
• Experience with Java, Reactor, and reactive programming.
• Experience with Spring WebFlux and Project Reactor.
• Experience modernizing legacy Java applications and transitioning from traditional Spring Framework architectures to reactive microservices.
• Experience mentoring engineering teams and providing architecture-level technical leadership.
• Experience working in principal-level engineering or architecture roles.