Inference Infrastructure Engineer, Serving

Elorian

$275K — $475K *
US-AnywhereRemote in Palo Alto, CA
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience with low-latency, high-throughput inference systems for large models
  • Strong knowledge of inference optimization techniques (e.g., quantization, batching)
  • Hands-on experience with serving frameworks (vLLM, TensorRT-LLM, Triton, SGLang)
  • Experience with multi-GPU/multi-node model parallelism
  • Strong systems programming skills (C++/CUDA a plus)
  • Experience with autoscaling and load balancing in ML services
  • Proven track record in GPU cost optimization at scale

Responsibilities

  • Build low-latency, high-throughput inference serving systems for multimodal models
  • Design techniques to enhance latency, throughput, and efficiency
  • Optimize codebase and GPU resources for maximum performance
  • Implement multi-GPU and multi-node parallelism in serving
  • Establish autoscaling and load balancing for ML services
  • Set standards for reliability, observability, and reproducibility
  • Collaborate with researchers to enhance inference for new architectures

Benefits

  • Health, dental, and vision benefits
  • Unlimited PTO
  • Paid parental leave
  • Relocation support as needed
Full Job Description
About the Role

We're looking for an infrastructure engineer to design, optimize, and scale the systems that serve our large multimodal models. Your work will make inference faster, more cost-effective, and more reliable, so our teams can focus on advancing model capabilities rather than managing bottlenecks.

Our focus is on performant, efficient inference, both to power real-world applications and to accelerate research. This role owns the infrastructure that ensures every deployment and evaluation runs smoothly at scale for our visual foundation models.

What You Will Do
  • Build low-latency, high-throughput inference serving systems for our large multimodal models
  • Design and implement techniques that improve latency, throughput, and efficiency, including quantization, batching, speculative decoding, and KV cache management
  • Optimize our codebase and GPU fleet to fully utilize hardware FLOPs, bandwidth, and memory
  • Implement multi-GPU and multi-node model parallelism for serving (tensor or pipeline parallel)
  • Build autoscaling and load balancing for production ML services
  • Establish standards for reliability, observability, and reproducibility across the inference stack
  • Collaborate with researchers to enable high-performance inference for novel architectures


Skills and Qualifications

Minimum qualifications:
  • 3+ years of experience building low-latency, high-throughput inference serving systems for large models
  • Strong knowledge of inference optimization techniques (quantization, batching, speculative decoding, KV cache management)
  • Hands-on experience with serving frameworks such as vLLM, TensorRT-LLM, Triton, or SGLang
  • Experience with multi-GPU/multi-node model parallelism for serving (tensor or pipeline parallel)
  • Strong systems programming skills; C++/CUDA a plus alongside Python
  • Experience with autoscaling and load balancing for production ML services
  • A track record of GPU cost optimization at scale

Preferred qualifications (strong candidates may have some, not all):
  • Experience serving multimodal (vision + language) models
  • Contributions to open-source ML or systems infrastructure projects (e.g., vLLM, SGLang, TensorRT-LLM, Triton)
  • A bias for action and comfort working across stacks and teams in an early-stage environment


Logistics
  • Location: This role is based on-site in Palo Alto, California.
  • Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $275,000 - $475,000 USD, plus equity and benefits.
  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.
  • Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

Similar Jobs

More Jobs at Elorian

More Enterprise Technology Jobs

Find similar Inference Infrastructure Engineer, Serving jobs: