About the RoleWe're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.
Essential Duties & Responsibilities- Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
- Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
- Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.
- Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.
- Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.
- Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
- Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.
- Work with model/system engineers to bring newly released models into production efficiently.
- Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.
- Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.
Qualifications- BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
- 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.
- Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.
- Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.
- Strong Python engineering skills.
- Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.
- Clear written and verbal communication skills to work effectively with a small, fully distributed team.
[
Preferred Qualifications (optional)]- Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc.
- Experience with MoE / long-context model deployment.
- Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.
- Experience with Nsight Systems / PyTorch Profiler.
- Familiarity with Kubernetes / production GPU serving.
- Previous startup experience.
Compensation & Benefits- Competitive salary with performance-based bonus and early-stage equity grant
- 100% employer-paid Health, Dental, and Vision coverage for you and your dependents
- 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals
- 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks
- Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)
- Daily lunch stipend
- Enterprise-level Claude & ChatGPT access with a generous token budget
- Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office
- Visa sponsorship and relocation assistance to one of our office hubs
- A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture