EPAM Systems

Principal AI Infrastructure & Accelerator Architect

EPAM Systems$175K — $210K *
US-AnywhereRemote in New York, NY
Enterprise Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or advanced degree in Computer Science, Electrical Engineering, or equivalent experience.
  • 15+ years of software engineering experience, with a focus on distributed systems, AI infrastructure, or HPC architecture.
  • 7+ years designing and developing complex software systems in C++ or Python.
  • 5+ years leading architecture and delivery of large-scale software products from inception to production.
  • Proven ability to architect performance-critical systems bridging hardware accelerators and high-level frameworks.

Responsibilities

  • Define multi-year technical roadmap for high-performance AI kernels and hardware-software co-design.
  • Mentor and scale a top-tier technical practice, establishing engineering standards.
  • Act as primary technical liaison with ML researchers and architecture teams to eliminate bottlenecks.
  • Architect foundational infrastructure, including benchmarking suites and automated tuning frameworks.
  • Track advancements in hardware architectures and compiler innovations for AI efficiency.

Benefits

  • Opportunity to shape the architectural vision impacting AI infrastructure at scale.
  • Work with cutting-edge hardware accelerators and advanced ML models.
  • Lead a prestigious Center of Excellence within a renowned technology company.
  • Collaborate closely with industry leaders and contribute to open-source ecosystems.
  • Access to continuous professional development and an elite technical team.
Full Job Description
Are you a visionary technical leader passionate about defining the future of AI infrastructure and squeezing every drop of performance out of advanced hardware accelerators at scale? We are seeking a Principal Architect to lead, shape, and execute our technical strategy for AI performance, optimization, and hardware-software co-design. In this elite, highly visible role, you will define the architectural vision for both AI training and serving infrastructure, delivering massive industry-wide impact. You will spearhead our Center of Excellence (CoE), scaling our practice and guiding the technical roadmap across next-generation Tensor Processing Units (TPUs), Graphics Processing Unit (GPU) fleets, state-of-the-art ML models, and advanced compiler toolchains. Your architectural decisions will directly enable cutting-edge AI research and large-scale production deployments across Google Cloud, major enterprise customers, and the broader open-source ecosystem. If you thrive on solving intractable performance bottlenecks and redefining what is physically and computationally possible in AI infrastructure, this is your platform. Req# [redacted] Responsibilities Define and drive the multi-year technical roadmap for high-performance AI kernels, custom operations, and hardware-software co-design targeting TPU and GPU architectures Scale and mentor a world-class technical practice, establishing architectural governance, engineering standards, and best practices across the organization Act as the principal technical liaison partnering with ML researchers, core framework architects (JAX, PyTorch), and compiler engineering teams (XLA, MLIR) to eliminate systemic bottlenecks and shape future hardware/software requirements Architect foundational infrastructure-including enterprise-grade benchmarking suites, automated autotuning frameworks, regression analysis pipelines, and comprehensive documentation-empowering the global developer community Anticipate industry shifts by tracking advancements in hardware architectures, emerging model topologies, and compiler innovations to unlock step-changes in AI training and inference efficiency Requirements Bachelor's degree in Computer Science, Electrical Engineering, or equivalent practical experience (Master's or Ph.D. preferred) 15+ years of software engineering experience, with 8+ years focused on distributed systems, AI infrastructure, or high-performance computing (HPC) architecture 7+ years of experience designing and developing complex software systems in C++ or Python 5+ years of experience leading the architecture, design, and delivery of large-scale software products, frameworks, or developer ecosystems from inception to production Proven track record of architecting performance-critical systems at the kernel level, bridging hardware accelerators and high-level software frameworks Nice to have Deep expertise in optimizing TPU/GPU execution, leveraging low-level kernel languages/abstractions such as Pallas, Mosaic, Triton, or CUDA Comprehensive knowledge of modern ML frameworks (JAX, PyTorch), attention mechanisms, Mixture of Experts (MoEs), model quantization, and low-precision arithmetic Advanced understanding of modern accelerator architectures, including heterogeneous compute, complex memory hierarchies, data movement optimization, and multi-node scale-out fabrics Deep familiarity with compiler principles, code generation, and modern toolchains such as MLIR, OpenXLA, and LLVM Demonstrated leadership in building and scaling developer infrastructure, widely adopted Open-Source Software (OSS) libraries, and extensible high-performance APIs Exceptional strategic communication and stakeholder management skills, with a history of influencing cross-functional engineering teams, researchers, and executive leadership

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Enterprise Technology Jobs

Find similar Principal AI Infrastructure & Accelerator Architect jobs: