About the RoleWe're looking for an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-effective, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.
Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research. This role owns the infrastructure that ensures every deployment and evaluation runs smoothly at scale for our visual foundation models.
What You Will Do- Build low-latency, high-throughput inference serving systems for our large multimodal models
- Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management
- Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory
- Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)
- Build autoscaling and load balancing for production ML services
- Establish standards for reliability, observability, and reproducibility across the inference stack
- Collaborate with researchers to enable high-performance inference for novel architectures
Skills and QualificationsMinimum qualifications:- 3+ years of experience building low-latency, high-throughput inference serving systems for large models
- Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)
- Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang
- Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)
- Strong systems programming skills; C++/CUDA a plus alongside Python
- Experience with autoscaling and load balancing for production ML services
- A track record of GPU cost optimization at scale
Preferred qualifications (strong candidates may have some, not all):- Experience serving multimodal (vision + language) models
- Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)
- A bias for action and comfort working across stacks and teams in an early-stage environment
Logistics- Location: This role is based on-site in Palo Alto, California.
- Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $275,000 - $475,000 USD, plus equity and benefits.
- Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
- Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.