AI Infrastructure Engineer

Netpreme

$120K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science or related field, or equivalent experience
  • 2+ years in LLM inference, ML systems, or performance engineering
  • Hands-on experience with vLLM and SGLang
  • Strong understanding of LLM inference fundamentals
  • Strong Python engineering skills
  • Familiarity with inference-serving concepts

Responsibilities

  • Deploy and optimize language and multimodal models using vLLM and SGLang
  • Design parallelism strategies across GPU systems
  • Build benchmarks to evaluate performance metrics
  • Analyze model architecture and its implications
  • Tune configurations for vLLM and SGLang
  • Profile and diagnose system bottlenecks
  • Collaborate with engineers to integrate new models into production

Benefits

  • 100% employer-paid Health, Dental, and Vision coverage
  • 401(k) match with immediate vesting
  • 100% employer-paid Life, Disability, and AD&D insurance
  • Generous PTO and company holidays
  • Daily lunch stipend
  • Access to AI resources like Claude & ChatGPT
  • Well-equipped office spaces with fitness amenities
  • Visa sponsorship and relocation assistance
Full Job Description
About the Role

We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.

Essential Duties & Responsibilities
  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.
  • Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.
  • Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.
  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
  • Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.
  • Work with model/system engineers to bring newly released models into production efficiently.
  • Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.
  • Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.

Qualifications
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.
  • Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.
  • Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.
  • Strong Python engineering skills.
  • Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.
  • Clear written and verbal communication skills to work effectively with a small, fully distributed team.

[Preferred Qualifications (optional)]
  • Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc.
  • Experience with MoE / long-context model deployment.
  • Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.
  • Experience with Nsight Systems / PyTorch Profiler.
  • Familiarity with Kubernetes / production GPU serving.
  • Previous startup experience.

Compensation & Benefits
  • Competitive salary with performance-based bonus and early-stage equity grant
  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents
  • 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals
  • 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks
  • Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)
  • Daily lunch stipend
  • Enterprise-level Claude & ChatGPT access with a generous token budget
  • Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office
  • Visa sponsorship and relocation assistance to one of our office hubs

Similar Jobs

More Jobs at Netpreme

More Information Technology Jobs

Find similar AI Infrastructure Engineer jobs: