Responsibilities
Your mission is to make inference so fast and cheap that evaluation never gates research.
- Build high-throughput inference systems for large-scale evaluation, backtesting, and scoring against historical physical observations
- Design and implement techniques that improve latency, throughput, and efficiency for real-time inference
- Optimize the inference stack to fully utilize hardware FLOPs, bandwidth, and memory
- Extend orchestration frameworks (e.g. Kubernetes, Ray, Slurm) for distributed inference and large-batch evaluation sweeps
- Establish standards for reliability, observability, and reproducibility across the inference stack, so every evaluation is trustworthy and repeatable
- Collaborate with researchers to enable high-performance inference for novel architectures as they emerge
What we're looking forWe value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
- Experience building or optimizing inference and serving systems for throughput and latency (e.g. TensorRT)
- Understanding of distributed compute, GPU parallelism, and hardware-aware optimization
- Deep familiarity with deep learning frameworks (e.g. PyTorch, JAX) and their underlying system architectures
- Strong engineering skills: performant, maintainable code and the ability to debug complex codebases
- Bonus: contributions to open-source inference or systems infrastructure (e.g. vLLM, SGLang, Triton)