AI/ML Training Performance Engineer

Blackrock Neurotech

$120K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience optimizing deep learning workloads for speed, memory efficiency, and cost reductions
  • Strong programming proficiency in Python and C++ or similar languages
  • Profound understanding of GPU architecture, execution, and performance optimization techniques
  • Hands-on expertise in developing and profiling GPU kernels using CUDA or HIP/ROCm
  • Familiarity with distributed training strategies and their impact on performance
  • Bachelor's degree in a relevant technical field or equivalent practical experience

Responsibilities

  • Own and improve performance of multi-GPU and multi-node deep learning training workloads
  • Profile training paths to identify and mitigate computational bottlenecks
  • Optimize tensor layouts and memory management for large model training
  • Develop and validate custom GPU kernels to enhance performance
  • Enhance Python training scripts and framework settings while maintaining training fidelity
  • Design scalable distributed training methodologies based on model topology
  • Coordinate with engineers to ensure data delivery matches training speed
  • Build reliable checkpoint and recovery systems for long-running experiments

Benefits

  • Opportunity for significant ownership over performance and scalability work
  • Collaborative work environment with a focus on high-quality outcomes
  • Emphasis on work-life balance and sustainable work practices
  • Potential for impact on high-consequence projects in AI/ML
  • Access to state-of-the-art GPU and cloud resources
Full Job Description
The Role

The ML Training Performance Engineer will own the efficiency and scalability of training models on GPU and cloud infrastructure. You will turn available compute into faster, more capable experiments as our model training scales in complexity and compute requirements.

As a hands-on individual contributor on a small research team, you will work across the training stack, from Python and model execution to GPU kernels, distributed communication, and runtime environments. Partnering with model researchers, data engineers, and infrastructure and IT teams, you will identify and implement performance improvements while preserving numerical correctness and scientific intent.

You will have significant ownership over how we measure, optimize, and scale training performance, establishing the baselines, tooling, and technical approaches that will support our AI/ML work as it grows.

What You'll Do
  • Own training performance across single-GPU, multi-GPU, and multi-node workloads, establishing reproducible baselines for throughput, memory use, utilization, time to target quality, and cost
  • Profile the full training path to distinguish compute, memory, communication, CPU, and I/O bottlenecks and prioritize changes with measurable end-to-end impact
  • Optimize tensor layouts, precision, memory allocation, activation checkpointing, operator fusion, and execution graphs to fit larger or longer-context models within available resources
  • Write, tune, and validate custom GPU kernels using CUDA, Triton, HIP/ROCm, or the appropriate platform tools when existing implementations limit performance
  • Improve Python training scripts, framework and compiler settings, batching, gradient accumulation, and optimizer execution while preserving intended training behavior
  • Design and tune distributed training strategies, including data, tensor, pipeline, or sharded parallelism, based on model structure, memory limits, and interconnect topology
  • Partner with model researchers on hardware-aware architecture and hyperparameter changes, measuring their effects on convergence, model quality, and compute requirements
  • Coordinate with the neural data infrastructure engineer on prefetching, pinned memory, host-to-device transfer, and I/O overlap so data delivery keeps pace with training
  • Work with infrastructure and IT on GPU selection, cloud instance configurations, networking, drivers, containers, scheduling, and capacity planning as training needs and compute capacity scale
  • Build robust checkpoint, restart, and recovery workflows and performance regression checks that keep long-running experiments reproducible and productive
  • Communicate benchmark evidence, numerical tradeoffs, scaling limits, and resource recommendations clearly to researchers and organizational stakeholders
What You Bring
  • Demonstrated experience improving the performance of substantial deep learning training workloads, with measured gains in speed, memory efficiency, or compute cost
  • A measurement-driven approach to performance optimization, using profiling and benchmarking to validate meaningful end-to-end improvements
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, or equivalent practical experience
  • Exceptional programming ability in Python and C++ or a comparable systems language, with strong debugging, testing, and performance analysis practices
  • Deep understanding of GPU execution, including memory hierarchies, memory coalescing, thread blocks, warps or wavefronts, occupancy, synchronization, and bandwidth limits
  • Hands-on experience developing and profiling GPU kernels with CUDA or HIP/ROCm, and the ability to diagnose correctness and performance at the hardware level
  • Deep knowledge of PyTorch or an equivalent framework, including automatic differentiation, computation graphs, tensor storage, compilation, and mixed-precision training
  • Strong understanding of deep learning architectures and the underlying computations that drive training performance
  • Experience with distributed training, collective communication, sharding, and the interaction between model partitioning and GPU interconnects
  • Strong understanding of numerical stability and the ability to validate gradients, convergence, and model quality after performance changes
  • Experience configuring and diagnosing Linux-based GPU environments, containers, cloud compute, and high-throughput storage or networking
  • Ability to collaborate closely with researchers and infrastructure teams and make clear tradeoffs between implementation effort, performance, reliability, and scientific value
  • Experience with Triton, compiler optimization, advanced GPU profiling tools, multiple accelerator generations, long-sequence or multimodal models, neural time series, or large-scale model training is a plus


Working Location:

This is an on-site role based at Blackrock Neurotech's headquarters in Salt Lake City, Utah. Occasional travel may be required.

How We Work

We are a small, experienced team working on consequential problems.

  • We take ownership of outcomes and follow through with clarity and accountability
  • We prioritize sustained, high-quality work over performative urgency
  • We value rigor, sound judgement and thoughtful decision-making
  • We collaborate deliberately: low ego, high trust and high context

This is a high-ownership role, but it is not an "always-on" one. We expect strong work and our people to have a life outside of it.

Similar Jobs

More Jobs at Blackrock Neurotech

More Information Technology Jobs

Find similar AI/ML Training Performance Engineer jobs: