The Model Optimization & ML Runtime team (within Smart Perception / Machine Learning) is responsible for maximizing the capability, efficiency, and hardware performance of Waymo's cutting-edge perception and foundation models. We bridge the gap between large-scale multi-task ML research (Vision Transformer backbones, 30+ perception heads, multi-sensor fusion) and real-time onboard vehicle deployment across custom automotive accelerators, developing core optimization frameworks (parameter-efficient fine-tuning, quantization, compilation, and quantization-aware training) that power the autonomous Waymo Driver.Waymo interns partner with leaders in the industry on projects that create impact to the company. We believe learning is a two-way street: applying your knowledge while providing you with opportunities to expand your skill-set. Interns are an important part of our culture and our recruiting pipeline. Join us at Waymo for a fun and rewarding internship!
You will:- Design, implement, and benchmark parameter-efficient fine-tuning (LoRA / QLoRA) modules in JAX/Flax for multi-task Vision Transformer backbones
- Develop dual-level distillation pipelines (intermediate feature matching and task-head logit distillation) to mitigate multi-task regressions during large-scale data scaling
- Collaborate with model optimization, quantization, and latency teams to validate static weight folding and low-precision quantization, ensuring zero latency overhead on onboard compute platforms
- Conduct extensive empirical ablations and evaluate perception metrics on large-scale autonomous driving datasets across diverse geographic domains
You have:- Currently pursuing a PhD or Master's in Computer Science, Electrical Engineering, Machine Learning, Robotics, or a related technical field
- Strong software engineering and deep learning development skills in Python and modern frameworks (JAX, Flax, PyTorch, or TensorFlow)
- Solid theoretical understanding and hands-on experience with deep learning foundation models, Transformer architectures, and multi-task learning
- Experience with model compression, parameter-efficient fine-tuning (e.g., LoRA, QLoRA, adapters), or quantization and knowledge distillation techniques
We prefer:- Publication record at top-tier computer vision or machine learning conferences (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR)
- Hands-on experience with model quantization (PTQ, QAT, INT8/INT4/MX4), low-precision numerics, or hardware-aware model optimization
- Experience training and scaling large vision backbones or multi-modal models on distributed accelerator clusters (TPUs / GPUs)
- Familiarity with autonomous driving perception tasks (3D object detection, semantics, tracking, or pedestrian intent prediction)
General Perks- Help solve challenging problems with a direct impact on the company
- Competitive compensation packages with a housing/relocation bonus (if applicable)
- Medical, dental, and vision insurance
- Fun intern events and networking opportunities
Onsite Perks- Free breakfast, lunch, dinner, and snacks
- Free access to Google shuttles
- Onsite gym
Note: This will be a hybrid onsite internship position. We will accept resumes on a rolling basis until the role is filled. To be in consideration for multiple roles, you will need to apply to each one individually - please apply to the top 3 roles you are interested in.
The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company's generous benefits programs, subject to eligibility requirements.
Hourly Masters Pay
$70-$70 USD
The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company's generous benefits programs, subject to eligibility requirements.
Hourly PhD Pay
$85-$85 USD