Full Job Description
You would collaborate with software engineers, AI researchers, and hardware specialists to develop high-performance solutions that meet the stringent requirements of autonomous driving applications. This is an exciting opportunity to work on next-generation transportation technology and make a meaningful impact on the future of mobility.
Key Responsibilities
• Optimize end-to-end GPU performance for real-time autonomous driving workloads, including sensor processing (e.g., camera, LiDAR) and neural network inference.
• Develop and optimize parallel computing algorithms and GPU-accelerated components using technologies such as CUDA.
• Collaborate with cross-functional teams to design and improve onboard GPU software architectures that meet the computational requirements of perception, planning, and control modules.
• Profile and analyze bottlenecks across GPU computation, memory access, data movement, synchronization, and CPU-GPU interaction.
• Debug and optimize GPU-based software to improve latency, throughput, resource utilization, and runtime stability on embedded platforms.
Qualifications:
Required:
• Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
• Strong knowledge of parallel computing principles, GPU architecture, memory hierarchy, and performance optimization techniques.
• Experience profiling GPU applications using tools such as NVIDIA Nsight Systems, Nsight Compute, or equivalent tools.
• Experience deploying or optimizing neural network inference workloads using technologies such as PyTorch, ONNX, and TensorRT.
• Experience with real-time embedded systems and handling large data streams from sensors (camera, LiDAR, radar).
• Strong proficiency in C/C++ and Python.
Preferred:
• 3+ years of experience in GPU programming and optimization (e.g., CUDA, OpenCL, Vulkan).
• Experience with NVIDIA Jetson Thor, NVIDIA DRIVE Thor, or similar embedded GPU platforms.
• Experience with model quantization, including FP8 and NVFP4.
• Experience managing concurrent GPU workloads and resource isolation using technologies such as NVIDIA Multi-Process Service (MPS), Multi-Instance GPU (MIG), or other related technologies.
• Experience with GPU-accelerated sensor data compression, including camera, LiDAR, or other onboard sensor data.