5-7 years of hands-on experience in machine learning optimizations
Proficient in CUDA programming and GPU architecture
Deep understanding of transformer and diffusion-based ML models
Experience with performance profiling and troubleshooting across hardware and software
Strong programming skills in Python and relevant libraries like PyTorch and NumPy
Responsibilities
Optimize GPU performance for both training and inference processes
Implement significant changes in ML architecture to achieve step-function improvements
Enhance both inference and training stacks for maximum efficiency
Analyze and address performance bottlenecks across CPU and network layers
Lead the integration of software and hardware modifications for optimization
Benefits
Innovative team environment that encourages experimentation and rapid advancement
Opportunities for professional development and continuous learning
Work on cutting-edge technology in a pioneering field
Collaborative culture focused on achieving ambitious goals
Full Job Description
About the Role
We internally call this team MBMB (More Big More Better). You will own optimizations on both the training and on-robot inference stacks. We are still in a regime of step-function, not incremental, gains.
You'll be responsible for:
Making GPUs go brrrrr
Implementing ML, hardware, and software changes that lead to step-function gains
Optimizing both the inference and training stacks
You might thrive in this role if you:
Are proficient and stay current with the latest ML techniques for training and inference optimizations in transformer and diffusion based architectures
Will chase ML optimizations anywhere: From the CUDA kernels, to ML architecture, to frontend or backend network bottlenecks, CPU bottlenecks, NVLink and comms, to torch, numpy, and Python inefficiencies.