You will:- Analyze and characterize internal feature representations of deep multimodal perception foundation models.
- Validate model performance on large-scale autonomous vehicle sensor datasets and simulation environments.
- Train, fine-tune, and evaluate deep neural networks to improve model performance.
- Inspect perception foundational model to derive signal on data quality.
- Implement or augment data pipeline to improve data quality
You have:- Currently enrolled in a PhD program in Computer Science, Robotics, Electrical Engineering, or a related quantitative field.
- Experience programming in Python.
- Practical experience training, fine-tuning, and evaluating deep learning models for computer vision or multimodal perception (e.g., using PyTorch, JAX, or TensorFlow).
- Solid understanding of modern neural architectures (e.g., Vision Transformers, multi modal sensor encoders).
We prefer:- Research experience or publications in multimodal deep learning models, or vision foundational models.
- Hands-on experience working with multi-camera and/or 3D LiDAR perception systems in robotics or autonomous driving.
- Experience evaluating large-scale perception systems under practical settings.
Note: This will be a hybrid onsite internship position. We will accept resumes on a rolling basis until the role is filled. To be in consideration for multiple roles, you will need to apply to each one individually - please apply to the top 3 roles you are interested in.
The expected hourly rate for this full-time position is listed below. Interns are also eligible to participate in the Company's generous benefits programs, subject to eligibility requirements.
Hourly PhD Pay
$85-$85 USD