Full Job Description
We are seeking a skilled Python Platform Engineer to operate and evolve the core infrastructure powering our enterprise knowledge base platform. As we transition from a Confluence-focused RAG chatbot into a highly advanced, multi-source agentic knowledge system, you will play a pivotal role in designing and scaling our AI backend. Today, our platform handles automated ingestion (from sources like internal wikis and code repositories) 1 chunking 1 pgvector storage 1 RAG retrieval 1 FastAPI serving. In this next phase, you will help us expand towards hybrid retrieval (combining vector, sparse, and graph search), multi-source ingestion pipelines, robust evaluation frameworks, and scalable agent infrastructure. Req.# [redacted] Responsibilities Architect and implement highly performant backend services using Python 3.11, FastAPI, Pydantic, SQLAlchemy async, and asyncpg Design critical retrieval trade-offs optimizing for quality, latency, operational cost, safety, and simplicity Build production-grade agent runtime capabilities including memory boundaries, tool sandboxing, granular permissions, and cost/budget controls Improve answer grounding, failure analysis, and citation enforcement (prioritizing robust production behavior over simple demo-only features) Create production-grade observability and feedback loops utilizing OpenTelemetry, Prometheus, Grafana, Docker, Helm, and GitHub Actions Partner closely with product and engineering teams to support multiple conversational surfaces through a unified knowledge platform Overhaul ingestion pipelines, manage AI workload profiles (handling latency, throughput, and failovers), and implement release workflows that validate complex AI behavior Requirements Strong, hands-on experience developing in Python within platform, automation, or infrastructure-heavy environments Proven experience building CLI tools utilizing Python, Golang, or Rust Deep experience working with LangGraph, LangChain, pgvector, and modern RAG/retrieval pipelines Experience designing and implementation-level knowledge of evaluation frameworks for LLM-backed systems (including regression detection and quality benchmarking) Strong experience with Docker, Helm, GitHub Actions, and Kubernetes-oriented container orchestrations Solid understanding of the operational characteristics, scaling bottlenecks, and cost profiles of embedding pipelines, vector search, and LLM providers Strong observability skills spanning metrics, tracing, alerting, dashboarding, and log analysis Experience managing ingestion, ETL, or large-scale content processing pipelines Nice to have Experience with specialized vector or graph infrastructure (e.g., Qdrant, Neo4j) Past experience supporting search platforms, RAG systems, or agent-based platforms in enterprise or highly regulated environments Familiarity with enterprise-grade tooling (e.g., Vault, Splunk, Artifactory, ECR) Comfort leveraging modern AI-assisted engineering tools (Copilot, etc.) to enhance your day-to-day coding workflow