EPAM Systems

Python Platform Engineer

EPAM Systems$120K — $150K *
US-AnywhereRemote in Canada
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Python development, particularly in platform or infrastructure contexts.
  • Proficiency in creating CLI tools using Python, Golang, or Rust.
  • Hands-on experience with LangGraph, LangChain, pgvector, and advanced RAG/retrieval systems.
  • Demonstrated ability in designing evaluation frameworks for LLM systems.
  • Familiarity with containerization technologies such as Docker, Helm, and Kubernetes.
  • Strong background in observability and monitoring practices for complex systems.

Responsibilities

  • Architect and maintain high-performance backend services utilizing Python 3.11 and FastAPI.
  • Optimize retrieval processes with a focus on quality and operational efficiency.
  • Develop robust agent runtime capabilities with strict memory and security controls.
  • Enhance answer grounding and failure analysis for improved production outcomes.
  • Implement effective observability measures using OpenTelemetry and Grafana.
  • Collaborate with cross-functional teams to support a unified knowledge platform.
  • Revamp ingestion pipelines, addressing performance metrics and managing AI workload profiles.

Benefits

  • Flexible work arrangements including remote options.
  • Professional development opportunities and continuous learning budget.
  • Health, dental, and vision insurance coverage.
  • Generous vacation policy to promote work-life balance.
  • Access to cutting-edge technology and tools for enhanced productivity.
Full Job Description
We are seeking a skilled Python Platform Engineer to operate and evolve the core infrastructure powering our enterprise knowledge base platform. As we transition from a Confluence-focused RAG chatbot into a highly advanced, multi-source agentic knowledge system, you will play a pivotal role in designing and scaling our AI backend. Today, our platform handles automated ingestion (from sources like internal wikis and code repositories) 1 chunking 1 pgvector storage 1 RAG retrieval 1 FastAPI serving. In this next phase, you will help us expand towards hybrid retrieval (combining vector, sparse, and graph search), multi-source ingestion pipelines, robust evaluation frameworks, and scalable agent infrastructure. Req.# [redacted] Responsibilities Architect and implement highly performant backend services using Python 3.11, FastAPI, Pydantic, SQLAlchemy async, and asyncpg Design critical retrieval trade-offs optimizing for quality, latency, operational cost, safety, and simplicity Build production-grade agent runtime capabilities including memory boundaries, tool sandboxing, granular permissions, and cost/budget controls Improve answer grounding, failure analysis, and citation enforcement (prioritizing robust production behavior over simple demo-only features) Create production-grade observability and feedback loops utilizing OpenTelemetry, Prometheus, Grafana, Docker, Helm, and GitHub Actions Partner closely with product and engineering teams to support multiple conversational surfaces through a unified knowledge platform Overhaul ingestion pipelines, manage AI workload profiles (handling latency, throughput, and failovers), and implement release workflows that validate complex AI behavior Requirements Strong, hands-on experience developing in Python within platform, automation, or infrastructure-heavy environments Proven experience building CLI tools utilizing Python, Golang, or Rust Deep experience working with LangGraph, LangChain, pgvector, and modern RAG/retrieval pipelines Experience designing and implementation-level knowledge of evaluation frameworks for LLM-backed systems (including regression detection and quality benchmarking) Strong experience with Docker, Helm, GitHub Actions, and Kubernetes-oriented container orchestrations Solid understanding of the operational characteristics, scaling bottlenecks, and cost profiles of embedding pipelines, vector search, and LLM providers Strong observability skills spanning metrics, tracing, alerting, dashboarding, and log analysis Experience managing ingestion, ETL, or large-scale content processing pipelines Nice to have Experience with specialized vector or graph infrastructure (e.g., Qdrant, Neo4j) Past experience supporting search platforms, RAG systems, or agent-based platforms in enterprise or highly regulated environments Familiarity with enterprise-grade tooling (e.g., Vault, Splunk, Artifactory, ECR) Comfort leveraging modern AI-assisted engineering tools (Copilot, etc.) to enhance your day-to-day coding workflow

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Information Technology Jobs

Find similar Python Platform Engineer jobs: