Rivian

Staff Software Engineer, ML Training and Inference Infrastructure

Rivian • $228K — $285K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent industry experience.
  • Deep expertise in PyTorch for machine learning applications.
  • Familiar with model training frameworks like PyTorch Lightning or Ray.
  • In-depth understanding of transformer architecture and optimizations for training and inference.
  • Experience in large-scale distributed model training.
  • Proven ability to profile and enhance model performance.

Responsibilities

  • Optimize Deep Learning training performance on NVIDIA GPU clusters.
  • Reduce latency for model inference and related processing on onboard systems.
  • Design, train, and deploy large deep learning models using extensive datasets.

Benefits

  • Comprehensive medical, prescription, dental, and vision insurance for employees and their families.
  • Coverage begins on day one of employment.
Full Job Description
Role Summary

As a Staff Software Engineer, ML training and inference infrastructure, you will be a member of the Perception team at Rivian, which develops advanced machine learning algorithms that directly impact safety critical self-driving features of our category defining vehicles.

We are looking for candidates with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training and inference of large autonomous driving models; and optimizing the training and inference performance.

Responsibilities

  • Optimize the performance of Deep Learning training workload on NVIDIA GPU systems on a large scale
  • Optimize the latency of model inference and model pre- and post-processing on onboard systems
  • Design, train, and deploy large deep learning models that can leverage the vast amount of labeled and unlabeled data

Qualifications

  • PhD in CS/CE/EE, or equivalent, in industry experience
  • Deep knowledge of PyTorch
  • Knowledge of model training framework (e.g. PyTorch Lightning, ray, etc.)
  • In-depth knowledge of transformer architecture and ways to accelerate the training and inference of transformer models
  • Experience of performing large scale distributed training of models
  • A track record of profiling models and doing detective work to improve model training and inference speed


Preferred Skill Requirements:
  • Experience with CUDA or Triton language for writing custom ops
  • Knowledge of Nvidia TensorRT
  • Knowledge of NCCL
  • Experience with edge computing systems
  • A track record of efficiently solving complex problems collaboratively on larger teams

Pay Disclosure

Salary Range for California Based Applicants: $228,000.00 - $285,000.00 (actual compensation will be determined based on experience, location, and other factors permitted by law).

Benefits Summary: Rivian provides robust medical/Rx, dental and vision insurance packages for full-time employees, their spouse or domestic partner, and children up to age 26. Coverage is effective on the first day of employment

Please note that we are currently not accepting applications from third party application services.

About Rivian

Rivian is an American automaker and automotive technology company. Founded in 2009, the company develops vehicles, products and services related to sustainable transportation. Rivian has raised over $10.5 billion since 2019, with investments from Amazon, Ford, and Cox Automotive. The company's first two vehicles, the R1T and R1S, are electric vehicles that are expected to be released in 2021. Rivian has also announced plans to produce electric delivery vans for Amazon. The company has received praise for its focus on sustainability and its commitment to using recycled materials in its vehicles.
Learn more about Rivian
Size
10,000 employees
Market Cap
$16.8 billion
Industry
Founded
2009
NASDAQ

Similar Jobs

More Jobs at Rivian

More Information Technology Jobs

Find similar Staff Software Engineer, ML Training and Inference Infrastructure jobs: