Full Job Description
We're looking for someone to architect and own the entire infrastructure pipeline behind that fleet: thousands of robots with embedded GPUs, communicating over wifi to inference clusters, streaming tens of petabytes of video to training clusters, with new model weights deployed every few minutes - all operating within strict real-time latency budgets.
This role requires someone who can hold an entire system in their head and optimize it end-to-end. You'll touch networking, storage, databases, embedded software, deployment systems, and GPU optimization - and you'll own the architecture decisions that tie them together.
You might be a good fit if you have:
- Designed and operated large-scale distributed systems from scratch
- Managed fleets of hundreds or thousands of computers
- Deep experience with GPU performance optimization, including writing custom CUDA kernels
- Worked on real-time embedded systems or robotics infrastructure
- Built petabyte-scale storage and database systems
- Experience with high-performance networking and video encoding pipelines
Nice to have:
- Rust and low-level performance optimization experience
- Experience taking infrastructure from prototype to production in a small team
We care much more about what you've built than any specific credential. We're a small, fast-moving team working together in person in San Francisco. If you're excited about architecting novel systems at unprecedented scale, we'd love to talk.