System Software Engineer, Robot Platform - GPU & Accelerated Compute

Sunday Inc

$120K — $160K *
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years of experience developing GPU systems software
  • Strong proficiency in CUDA and a systems language such as C++, C, or Rust
  • Solid understanding of GPU architecture and workload management
  • Hands-on experience with the CUDA ecosystem: runtime API, Graphs, IPC
  • Familiarity with GPU sharing mechanisms like MPS and MIG
  • Experience with GPU profiling tools such as Nsight Systems and Compute
  • Solid Linux fundamentals, including scheduling and performance tuning

Responsibilities

  • Own and enhance the accelerated compute layer of the robot platform
  • Reduce GPU kernel launch overheads for efficient model execution
  • Arbitrate GPU access to ensure predictable latency across applications
  • Drive low-latency transfer of camera frames to GPU memory
  • Build efficient data movement between CPU and GPU, including zero-copy paths
  • Design synchronization patterns to minimize stalls during inference
  • Collaborate with cross-functional teams to optimize GPU resource utilization

Benefits

  • Collaborative work environment with cross-functional teams
  • Opportunity to work on cutting-edge robotics technology
  • Access to advanced GPU and accelerated computing resources
  • Career growth and development opportunities
  • Potential contributions to industry-leading CUDA libraries
Full Job Description
What to Expect

The Robot Platform team builds the foundational systems that every part of our robot perception, ML, controls and behavior runs on, and the developer infrastructure that lets us build, ship, and update that software quickly and safely on every robot in the fleet.

As a System Software Engineer on Robot Platform focused on GPU and accelerated compute, you'll own how every accelerated workload on the robot from model inference, SLAM/perception, and more gets data, gets scheduled and runs efficiently on shared compute. You'll work alongside teammates who own the runtime and our build and delivery infrastructure, and you'll partner cross-functionally with ML, SLAM/Perception, Controls and Hardware teams to ensure the GPU is a first-class, well-utilized resource that meets the latency and throughput requirements of a real-time robotic system operating in the home.

What You'll Do

You'll own and contribute to the accelerated compute layer of the robot platform, including:
  • Efficient model execution and switching: Reduce gpu kernel launch overheads and make swapping between models on the same device fast and predictable
  • GPU scheduling and time-slicing: Arbitrate GPU access across concurrent users (model inference, SLAM, and other robotics applications) with predictable latency
  • Camera pipeline: Drive low-latency transfer of camera frames into GPU memory, integrating with HW accelerate encode/decode (NVDEC/NVENC) where appropriate
  • CPU 1212 GPU data transfer: Build efficient, low-overhead data movement between host and device, including pinned memory, zero-copy paths, and asynchronous transfer patterns
  • CPU/GPU synchronization: Design synchronization primitives and patterns that minimize stalls and keep inference pipelines full


What You'll Bring
  • 2+ years of experience developing gpu systems software
  • Strong proficiency in CUDA and a systems language such as C++, C, or Rust
  • Solid understanding of GPU architecture, GPU workloads, and the tradeoffs involved in time-slicing and sharing the device across users
  • Hands-on experience with the CUDA ecosystem: CUDA runtime API, CUDA Graphs, and CUDA IPC
  • Familiarity with GPU sharing mechanisms such as MPS and MIG
  • Experience with GPU profiling tools such as Nsight Systems and Nsight Compute
  • Solid Linux fundamentals: scheduling, IPC, memory management, and performance tuning


Nice to Have
  • Contributions to CUDA libraries or other GPU programming libraries
  • Experience with camera pipeline integration and NVDEC/NVENC
  • Experience optimizing model inference on embedded GPU platforms (e.g., Jetson)
  • Experience with observability and tracing for GPU-accelerated workloads


Similar Jobs

More Jobs at Sunday Inc

  • Product Operations
    $90K — $130K *
    Redwood City, CA 94061 (San Mateo County)
    Consumer Technology
    In-Person
  • Triage Associate
    $70K — $95K *
    Redwood City, CA 94061 (San Mateo County)
    Technical Services
    In-Person
  • Eval Ops Program Manager
    $100K — $130K *
    Redwood City, CA 94061 (San Mateo County)
    Information Technology
    In-Person
  • Data Annotation Lead
    $120K — $150K *
    Redwood City, CA 94061 (San Mateo County)
    Information Technology
    In-Person
  • R&D Engineering Technician
    $70K — $95K *
    Redwood City, CA 94061 (San Mateo County)
    Technical Services
    In-Person

More Technical Services Jobs

Find similar System Software Engineer, Robot Platform - GPU & Accelerated Compute jobs: