SummaryYou will bridge the gap between quantitative research and high-performance computing, building and optimizing the systems used to train machine learning models at scale. You will focus on accelerating the end-to-end training lifecycle-from data ingestion and distributed execution to kernel performance and hardware utilization-enabling researchers to iterate more quickly across increasingly complex models and datasets.
Responsibilities- Training Performance and Benchmarking
- Benchmark model-training workloads across CPUs, GPUs, and other accelerator platforms to identify bottlenecks and guide Tower's compute infrastructure decisions.
- Develop performance models and standardized benchmarks for measuring throughput, utilization, scalability, and time to convergence.
- Distributed Training Optimization
- Design and optimize distributed training strategies, including data, tensor, pipeline, and model parallelism.
- Improve communication efficiency across multi-GPU and multi-node environments by optimizing collective operations, topology awareness, and computation-communication overlap.
- End-to-End Training Efficiency
- Analyze and improve the full training pipeline, including data loading, preprocessing, memory management, forward and backward passes, optimizer execution, checkpointing, and experiment recovery.
- Identify bottlenecks across compute, memory, storage, networking, and interconnects to increase accelerator utilization and researcher productivity.
- GPU Kernel and Framework Development
- Develop and optimize GPU kernels and performance-critical framework components for quantitative machine learning workloads.
- Integrate specialized libraries, compilers, and execution techniques to improve throughput, memory efficiency, and numerical performance.
- Model and Numerical Optimization
- Apply techniques such as mixed-precision training, gradient accumulation, activation checkpointing, operator fusion, and memory-efficient optimizers.
- Evaluate tradeoffs among training speed, numerical stability, reproducibility, model quality, and infrastructure cost.
- Training Infrastructure
- Partner with HPC and infrastructure teams to optimize workload scheduling, resource allocation, observability, fault tolerance, and reproducibility across shared compute environments.
- Help define the architecture and tooling required to support large-scale experimentation across on-premises and cloud-based infrastructure.
- Cross-Functional Collaboration
- Work closely with ML Researchers, Quantitative Researchers, HPC Engineers, Systems Engineers, and hardware specialists to translate research requirements into highly efficient training systems.
Qualifications- 3+ years of experience optimizing machine learning training workloads in high-performance, distributed, or large-scale computing environments.
- Deep knowledge of machine learning frameworks such as PyTorch or JAX, including their execution models, compilation paths, autograd systems, and distributed-training capabilities.
- Strong programming skills in Python and C++, with experience developing or optimizing performance-critical systems.
- Proven experience with GPU kernel development and optimization using technologies such as CUDA, Triton, CUTLASS, cuBLAS, cuDNN, or related libraries.
- Strong understanding of GPU architecture, including streaming multiprocessor execution, warp scheduling, tensor cores, and the memory hierarchy from registers through HBM.
- Experience with distributed-training technologies and communication libraries such as NCCL, FSDP, DeepSpeed, Megatron-LM, XLA, or equivalent systems.
- Proficiency with performance-analysis tools such as Nsight Systems, Nsight Compute, PyTorch Profiler, or comparable tracing and profiling platforms.
- Understanding of high-performance networking, storage, and accelerator interconnects, including technologies such as InfiniBand, RDMA, NVLink, or NVSwitch.
- Demonstrated ability to benchmark heterogeneous compute platforms and make rigorous, data-driven recommendations about performance, scalability, and cost.
Preferred Qualifications- Experience optimizing training workloads for transformer-based, time-series, reinforcement-learning, or other computationally intensive models.
- Experience with cluster orchestration and scheduling technologies such as Kubernetes, Slurm, Ray, or similar platforms.
- Familiarity with fault-tolerant distributed training, large-scale checkpointing, experiment reproducibility, and GPU-cluster observability.
- Practical experience with specialized accelerators, custom hardware, or compiler technologies for machine learning.
- Prior experience in financial trading is not required.
Anticipated New York annual base salary of $200,000, plus eligible for discretionary bonus.
BenefitsTower's headquarters are in the historic Equitable Building, right in the heart of NYC's Financial District and our impact is global, with over a dozen offices around the world.
At Tower, we believe work should be both challenging and enjoyable. That is why we foster a culture where smart, driven people thrive - without the egos. Our open concept workplace, casual dress code, and well-stocked kitchens reflect the value we place on a friendly, collaborative environment where everyone is respected, and great ideas win.
Our benefits include:
- Generous paid time off policies
- Savings plans and other financial wellness tools available in each region
- Hybrid working opportunities
- Free breakfast, lunch and snacks daily
- In-office wellness experiences and reimbursement for select wellness expenses (e.g., gym, personal training and more)
- Volunteer opportunities and charitable giving
- Social events, happy hours, treats and celebrations throughout the year
- Workshops and continuous learning opportunities
At Tower, you'll find a collaborative and welcoming culture, a diverse team and a workplace that values both performance and enjoyment. No unnecessary hierarchy. No ego. Just great people doing great work - together.