Senior Principal AI Engineer

Cerence Inc.$150K — $180K *
US-AnywhereRemote in United States
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience with distributed systems or ML systems
  • Proven ability to manage large-scale workloads on GPU clusters
  • Experience with PyTorch distributed training in production environments
  • In-depth knowledge of parallelism strategies including data, tensor, and pipeline parallelism
  • Understanding of GPU communication and networking at a low level
  • Familiarity with orchestration tools like Slurm, Kubernetes, Ray, RunAI
  • Competence with memory optimization techniques such as activation checkpointing and ZeRO strategies.

Responsibilities

  • Design and operate distributed training systems for various neural network models across GPU clusters
  • Optimize execution of multi-node, multi-GPU setups to enhance throughput
  • Diagnose and resolve bottlenecks in compute, memory, and networking
  • Enhance training stability and fault tolerance for large scale deployments
  • Collaborate with research and applied ML teams to implement robust training pipelines
  • Build and optimize orchestration frameworks for GPU clusters
  • Minimize networking bottlenecks to improve overall training efficiency.

Benefits

  • Opportunity to work on cutting-edge AI technology
  • Direct impact on large model training and research efficiency
  • Collaborative work environment with access to ML experts
  • Focus on developing scalable, reliable AI infrastructure
  • Potential for professional growth in a rapidly evolving field.
Full Job Description
A Moving Experience.

What You Will Work On

  • Design and operate distributed training systems for large neural networks (autoregressive, diffusion, State Space Models etc.) across GPU clusters


  • Optimise multi-node, multi-GPU execution to maximize throughput and utilization


  • Diagnose & resolve bottlenecks across compute, memory, and network


  • Improve training stability and fault tolerance at scale


  • Partner with research and applied ML teams to productionize large-model training pipelines


Core Responsibilities

  • Distributed Training Infrastructure


  • Build and optimize GPU cluster orchestration using:


  • Slurm


  • Kubernetes


  • Ray


  • RunAI


  • Ensure efficient scheduling, isolation, and fairness across training workloads


  • Communication & Networking


  • Optimize and debug distributed communication using:


  • NCCL


  • RDMA


  • InfiniBand


  • NVLink


  • Minimize networking bottlenecks that dominate end-to-end training time


  • Training Frameworks


  • Scale large-model training using:


  • PyTorch Distributed


  • Megatron-LM


  • DeepSpeed


  • Own multi-node launch configurations, failure recovery, and performance tuning


  • Memory & Performance Optimization


  • Apply advanced memory optimization techniques:


  • Activation checkpointing


  • ZeRO (Stage 1-3) and offload strategies


  • Balance compute, memory, and communication to push model size and batch scale


What Success Looks Like

  • GPU utilization consistently stays high (>80-90%)


  • Training scales cleanly from single node to dozens or hundreds of GPUs


  • Communication overhead is minimized and predictable


  • Large training jobs run stably for days or weeks without failure


  • New models can be trained faster, larger, and more reliably than before


Required Experience & Skills

  • Strongly Required


  • Deep hands-on experience with distributed systems or ML systems


  • Experience running large-scale workloads on GPU clusters


  • Production experience with PyTorch distributed training


  • Strong understanding of parallelism strategies (data, tensor, pipeline parallelism)


  • Low-level understanding of GPU communication and networking


  • Critical Technical Skills


  • GPU orchestration: Slurm, Kubernetes, Ray, RunAI


  • Communication libraries: NCCL, RDMA, InfiniBand, NVLink


  • Training frameworks: PyTorch Distributed, Megatron-LM, DeepSpeed


  • Memory optimisation: activation checkpointing, ZeRO offload techniques


Common Problems You'll Be Solving

  • Many teams fail at scale because:


  • GPU utilization is low despite large clusters


  • Networking and communication dominate training time


  • Training jobs crash or become unstable at large scale


  • You will be explicitly focused on eliminating these failure modes.


Ideal Background

  • This role is a strong fit for individuals who have worked as:


  • ML Systems Engineer


  • Distributed Systems Engineer


  • AI Infrastructure Engineer


  • HPC Engineer transitioning into ML


  • Experience working with large language models or foundation models is a strong plus, but deep systems expertise is valued over pure model architecture experience.


Why This Role Matters

Without robust distributed training infrastructure, progress on large models stalls. This role directly enables:

  • Larger models


  • Faster iteration cycles


  • More reliable research-to-production pipelines


You will be building the foundation that makes large-scale AI possible.

About Cerence Inc.

Cerence Inc. is a software company that specializes in voice recognition and natural language understanding technology. The company was spun off from Nuance Communications in 2019 and is headquartered in Newton, Massachusetts. Cerence's software is used in a variety of applications, including automotive infotainment systems, smart speakers, and virtual assistants. The company's clients include many of the world's leading automakers, as well as companies in the consumer electronics and mobile device industries. Cerence has received several awards for its technology, including the 2020 CES Innovation Award for its Cerence Drive platform.
Learn more about Cerence Inc.
Size
1,200 employees
Market Cap
$726.1 million
Industry
Net Income
$12.7 million
5 Year Trend
+6%
Revenue
$347.1 million

Similar Jobs

More Jobs at Cerence Inc.

More Technical Services Jobs

Find similar Senior Principal AI Engineer jobs: