EPAM Systems

Lead AI Agent Engineer

EPAM Systems • $150K — $180K *
US-AnywhereRemote in Georgia, US
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in delivering and operating production software, including LLM-based agents.
  • Strong programming skills in at least one general-purpose language with practical agent development experience.
  • Experience with cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI).
  • Proven track record of technical leadership and team delivery.
  • Ability to work across language boundaries and guide technology choices.

Responsibilities

  • Lead the design of AI agent solutions, focusing on execution, memory, and observability.
  • Translate customer needs into clear technical requirements and justify chosen solutions.
  • Guide the team in using cloud services for AI hosting and establish implementation patterns.
  • Design and implement an LLM gateway and routing layer for model management and observability.
  • Set integration standards for security workflows and define permissions and auditability.
  • Establish a balanced testing strategy covering software correctness and operational performance.
  • Plan and prioritize technical delivery, managing dependencies and resolving impediments.

Benefits

  • Opportunity to lead innovative AI projects in a hands-on technical role.
  • Collaborative environment with a focus on mentoring and team development.
  • Access to cutting-edge cloud technologies and AI tools.
  • Flexibility to work with engineers from diverse programming backgrounds.
  • Engagement in complex problem-solving and troubleshooting activities.
Full Job Description
Lead the technical design and delivery of production AI agent solutions. This is a hands-on technical leadership role combining architecture, engineering, and team delivery. You will work with customers and stakeholders to define the right solution, guide engineers through implementation, and take responsibility for technical quality and operational readiness across the team's work. We welcome engineers from different programming backgrounds. We value depth in your existing language and the ability to learn the languages and tools needed for the role. Responsibilities Lead solution design across agent execution, memory, identity, tool integration, and observability Turn customer needs into clear requirements, compare alternatives, and explain the business and technical reasons for the chosen approach Guide the team's use of cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI) Establish implementation patterns and validate platform choices against reliability, security, cost, and operational requirements Design and guide implementation of an LLM gateway & routing layer covering model tiering, quota/cost control, and observability (e.g., LiteLLM, APIM AI Gateway) Design and review memory, context, and state-management approaches for long-running workflows, including recovery, data retention, access boundaries, and isolation between users or engagements Set integration standards for MCP servers and tools used in authorized security workflows, including reconnaissance, scanners, controlled exploit tooling, and internal services Define permissions, approval requirements, auditability, and containment of failures Establish a balanced testing and evaluation strategy covering software correctness, agent behavior, security boundaries, and operational performance Define release criteria and ensure the team can demonstrate that they are met Guide production readiness and incident response, including observability, cost controls, staged delivery, rollback, and recovery Use operational evidence to prioritize improvements and technical debt Plan and prioritize technical delivery with product, security, and engineering stakeholders Break work into achievable milestones, delegate ownership, manage dependencies, and resolve technical and delivery impediments Personally build working prototypes and production-grade slices of the solution, taking them through implementation, testing, and deployment Lead complex troubleshooting and develop other engineers through design reviews, coaching, and constructive feedback Improve engineering processes, contribute to technical hiring, and communicate progress, risks, alternatives, and decisions clearly to customers and the team Requirements Substantial experience delivering and operating production software, typically five or more years, including at least one year shipping LLM-based agents to production You also bring a demonstrated record of technical leadership and team delivery Strong programming skills and production depth in at least one general-purpose language, with practical agent development experience You can work across language boundaries, guide technology choices, and help the team adopt unfamiliar tools You personally deliver working prototypes and production-grade implementations Experience designing system components or complete solutions, evaluating architectural alternatives, and making decisions against business needs and non-functional requirements Deep experience with cloud managed services for AI hosting (e.g., Bedrock, Azure AI Foundry, Vertex AI), including runtime, state, identity, integration, and observability concerns, and evidence of leading a team through adoption of an unfamiliar platform Practical expertise in MCP, tool orchestration, context and memory management, and agent evaluation, supported by a strong understanding of distributed-system failures and security boundaries Experience establishing or improving testing, CI/CD, observability, and engineering practices You can explain how these practices improved quality, delivery, or operational outcomes Experience leading technical discussions with customers, presenting alternatives, and aligning stakeholders on scope, priorities, risks, and trade-offs The ability to plan team delivery, delegate effectively, resolve disagreements and impediments, mentor engineers, and conduct technical interviews Sound judgment when adopting AI development tools: clear expectations for data access, permissions, review, and verification, with outcomes assessed through evidence English proficiency at B2 level or higher Nice to have Experience leading engineering work in application security, penetration testing, red-teaming, or security automation within an authorized scope Experience designing platforms for sandboxed code execution, browser agents, or autonomous tool orchestration Experience establishing reusable agent components, evaluation practices, or operational standards used by multiple teams Experience taking a new product or technical capability from initial design through launch and ongoing operation

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Information Technology Jobs

Find similar Lead AI Agent Engineer jobs: