EPAM Systems

Senior Manager, Technology Consulting

EPAM Systems$150K — $180K *
US-AnywhereRemote in United States
Enterprise Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor’s degree or equivalent experience
  • 12+ years in the industry with 5 years in software development in C++ or Python
  • 3 years testing or launching software products
  • 1 year in software design and architecture
  • Experience optimizing TPU/GPU code in low-level languages
  • Knowledge of ML frameworks like JAX and PyTorch
  • Proven background in building developer infrastructure and OSS libraries.

Responsibilities

  • Design and optimize high-performance kernels for ML operations using TPU and GPU architectures
  • Architect benchmarking suites and performance analysis tools to enhance developer interaction
  • Track hardware architecture and compiler advancements for optimization opportunities
  • Engage with ML researchers and framework developers to enhance adoption and address bottlenecks
  • Create documentation to support the adoption of custom kernels in OSS libraries

Benefits

  • Opportunity to work with cutting-edge AI and hardware technologies
  • Impactful role shaping the future of AI at a leading tech company
  • Access to advanced toolchains for ML optimization
  • Engagement with a collaborative and innovative team
  • Enhanced exposure to a broad open-source ecosystem
Full Job Description
Senior Manager, Technology Consulting Are you passionate about squeezing every last drop of performance out of advanced hardware accelerators? Step into a role where you will shape the future of AI. In this role, you will drive the performance and optimization of both training and serving, delivering massive impact for customers. We are seeking a talented leader to build out a center of excellence and scale this practice. You will have exposure to the newest Tensor Processing Unit (TPU) and Graphics Processing Unit (GPU) hardware, the latest ML models, and the advanced toolchains that bridge them. Your work will directly enable AI research and production deployments across Google Cloud and the broader open-source ecosystem. You will address complex technical issues that directly impact the efficiency and scalability of AI across the industry. The AI and Infrastructure team is redefining what's possible. Req# [redacted] Responsibilities Design and optimize high-performance kernels (using languages like Pallas, Mosaic, and Triton) targeting Tensor Processing Unit (TPU) and Graphics Processing Unit (GPU) architectures for critical Machine Learning (ML) operations, redefining what's possible from massive training runs to high-speed inference Architect infrastructure such as benchmarking suites, autotuning frameworks, performance analysis tools, regression testing, and documentation, transforming how the developer community interacts with increasingly critical custom kernels in key Open-Source Software (OSS) libraries Track the latest advancements in hardware architectures, compiler technologies, and AI models to identify new opportunities for performance optimization through custom kernels Engage with ML researchers, framework developers (Just After eXecution (JAX), PyTorch), and compiler engineers (Accelerated Linear Algebra (XLA)) to enhance adoption, identify new requirements, and address bottlenecks by providing appropriate solutions Requirements Bachelor’s degree or equivalent practical experience Overall 12+ years of industry experience; 5 years of experience with software development in C++ or Python 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture Experience with performance optimization at the kernel level Experience optimizing TPU/GPU code, using low-level kernel languages like Pallas, Compute Unified Device Architecture (CUDA), or Triton Knowledge of ML Frameworks (JAX/PyTorch), common operations like attention and Mixture of Experts (MoEs), including model optimization and low-precision formats Understanding of modern accelerators (e.g., data movement, pipelining, heterogeneous compute, and scale-out) Understanding of compiler principles (optimization, code generation) and toolchains such as MLIR, OpenXLA Demonstrate a record of building developer infrastructure, including Open-Source Software (OSS) libraries, flexible high-performance APIs, and easy-to-consume documentation to empower the community Excellent investigative and problem-solving capabilities with communication skills across cross-functional teams

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Enterprise Technology Jobs

Find similar Senior Manager, Technology Consulting jobs: