Role SummaryAs a Staff Software Engineer for ML Optimization and Hardware Acceleration, you will be a lead member of the Autonomy team at Rivian. You will develop and optimize advanced machine learning algorithms that directly impact the safety-critical self-driving features of our category-defining vehicles. This role focuses on the intersection of cutting-edge model architectures including Transformers, LLMs, VLMs, LDMs and high-performance hardware execution. You will bridge the gap between theoretical ML research and real- time embedded deployment, ensuring our autonomy stack remains both state-of-the-art and ultra-efficient.
Responsibilities- Model Optimization: Develop and deploy ultra-low latency Deep Learning and Machine Learning algorithms specifically tailored for Rivian ADAS and Autonomy use cases.
- Hardware-Aware Design: Research and implement hardware-aware optimization strategies, including Post-Training Quantization (PTQ), Quantization-Aware Training (QAT), kernel fusion, and model distillation to maximize throughput on embedded platforms.
- Performance Profiling: Utilize and automate deep-dive profiling tools (e.g., Torch Profile, NVIDIA Nsight) to identify bottlenecks and ensure performance consistency across weekly evaluation runs.
- Cross-Functional Collaboration: Partner with low-level software and hardware architecture teams to characterize in-house ML models on embedded platforms, optimizing them within strict compute and memory constraints.
- Architectural Reasoning: Apply a deep understanding of GPU architectures to optimize models across significantly different hardware targets, ensuring scalability across the Rivian fleet.
- Workflow and Infrastructure Engineering: Design and build automated pipelines for regular model profiling across diverse architectures to enhance organization-wide insight into execution bottlenecks.
Qualifications- Education/Experience: MS (+3 years of experience in deep learning, heterogeneous computing, and ML accelerators) or Ph.D. in Computer Science, Electrical Engineering, or a related field.
- Core ML Expertise: Deep understanding of modern model architectures, including Transformers, LLMs, VLMs and LDMs.
- Optimization Skills: Proven experience in model compression techniques: knowledge distillation, pruning, and quantization (PTQ/QAT).
- Hardware Knowledge: In-depth understanding of GPU architecture and the ability to optimize for diverse hardware specifications.
- Technical Toolset:
• Proficiency in Python and deep knowledge of PyTorch or TensorFlow.
• Hands-on experience with TensorRT, AIMET, ONNX runtimes.
• Experience with low-level programming (CUDA kernels, C++, or BLAS subroutines) for inference logic.
• Experience with profiling tools like torch profiler and nvidia nsight. - Leadership: Strong team player with excellent communication skills to drive complex, cross-functional efforts in a fast-paced environment.
How to distinguish yourself:
• A strong track record of publications in top-tier venues such as MLSys, ICML, NeurIPS, or ISCA.
• Significant and direct industry experience in a related domain.
• Active participation and contributions to relevant open-source projects.
• Public demonstrations of expertise, including technical talks, presentations, or live demos.
Pay DisclosureSalary Range for California Based Applicants: $228,000 - $285,000 (actual compensation will be determined based on experience, location, and other factors permitted by law).
Benefits Summary: Rivian provides robust medical/Rx, dental and vision insurance packages for full-time employees, their spouse or domestic partner, and children up to age 26. Coverage is effective on the first day of employment, and Rivian overs most of the premiums.