AI Infrastructure Engineer

Netpreme

$130K — $155K *
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Science, Computer Engineering, or equivalent experience.
  • 2+ years of experience in LLM inference, ML systems, GPU systems, or performance engineering.
  • Hands-on experience deploying and optimizing vLLM and/or SGLang.
  • Strong understanding of LLM inference fundamentals and distributed GPU execution.
  • Strong Python engineering skills.
  • Working knowledge of inference-serving concepts like continuous batching and KV cache handling.
  • Excellent communication skills for effective collaboration.

Responsibilities

  • Deploy and optimize language and multimodal models with vLLM and SGLang.
  • Design and evaluate various parallelism strategies across GPU systems.
  • Build and analyze reproducible benchmarks for model performance metrics.
  • Diagnose and address performance bottlenecks across GPU resources.
  • Tune configurations related to GPU utilization and throughput.
  • Collaborate with engineers to efficiently deploy new models into production.
  • Contribute to the definition of benchmarks for evolving workloads.

Benefits

  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents.
  • 401(k) match with immediate vesting and financial advisor access.
  • 100% employer-paid Life, Disability, and AD&D insurance, with wellness perks.
  • 20 vacation days and 15 company holidays, including floating days.
  • Daily lunch stipend and generous access to AI tools like Claude & ChatGPT.
  • Well-equipped offices with on-site amenities like fitness centers and parking.
  • Visa sponsorship and relocation assistance where applicable.
Full Job Description
About the Role

We're looking for an AI Infrastructure Engineer to build and operate the serving infrastructure You will work hands-on with vLLM and SGLang on Kubernetes. This is a foundational infrastructure role on a small, high-autonomy team.

Essential Duties & Responsibilities
  • Deploy and optimize large language and multimodal models using vLLM and SGLang or other inference engines.
  • Design and evaluate TP/EP/DP/PP and hybrid parallelism strategies across GPU systems.
  • Build reproducible benchmarks to evaluate TTFT, TPOT, throughput, concurrency scaling, GPU utilization, and memory utilization.
  • Analyze model architecture and its serving implications, including attention, KV cache, MoE, long context, and speculative decoding.
  • Tune vLLM and SGLang configurations such as continuous batching, max batched tokens, chunked prefill, prefix caching, KV-cache precision/capacity, speculative decoding, CUDA Graphs, and P/D disaggregation.
  • Profile and diagnose bottlenecks across GPU compute, memory, communication, scheduling, and serving runtime.
  • Compare deployment configurations and identify production operating points balancing latency, throughput, capacity, and stability.
  • Work with model/system engineers to bring newly released models into production efficiently.
  • Collaborate closely with our hardware/systems team (direct access to CTO-level technical leadership on a small team) to translate performance requirements into backend architecture decisions.
  • Contribute to defining next-generation benchmarks and service requirements as workloads evolve - multi-turn coding, agentic pipelines, RAG, and other long-context use cases.

Qualifications
  • BS, MS, or PhD in Computer Science, Computer Engineering, or a related field, or equivalent experience.
  • 2+ years of relevant experience in LLM inference, ML systems, GPU systems, or performance engineering.
  • Must have: hands-on experience deploying and performance-tuning vLLM and/or SGLang.
  • Strong understanding of LLM inference fundamentals, including prefill vs. decode, batching, KV cache, latency/throughput trade-offs, and distributed GPU execution.
  • Strong Python engineering skills.
  • Working knowledge of inference-serving concepts: continuous batching, KV cache handling, quantization, and serving SLAs.
  • Clear written and verbal communication skills to work effectively with a small, fully distributed team.

[Preferred Qualifications (optional)]
  • Contributions to vLLM, SGLang, FlashInfer, TensorRT-LLM, etc.
  • Experience with MoE / long-context model deployment.
  • Experience with speculative decoding, prefix caching, P/D disaggregation, attention/KV optimization.
  • Experience with Nsight Systems / PyTorch Profiler.
  • Familiarity with Kubernetes / production GPU serving.
  • Previous startup experience.

Compensation & Benefits
  • Competitive salary with performance-based bonus and early-stage equity grant
  • 100% employer-paid Health, Dental, and Vision coverage for you and your dependents
  • 401(k) match with immediate vesting, and access to financial advisors to help you reach your financial goals
  • 100% employer-paid Life, Disability, and AD&D insurance, plus a fitness stipend and wellness & mental health perks
  • Generous PTO: 20 vacation days, 15 company holidays (including 3 floating days of your choosing)
  • Daily lunch stipend
  • Enterprise-level Claude & ChatGPT access with a generous token budget
  • Well-equipped, sunny offices in Santa Clara, CA & Cambridge, MA with on-site parking and EV charging; on-site fitness center in Santa Clara; gym discounts near our Cambridge office
  • Visa sponsorship and relocation assistance to one of our office hubs
  • A collaborative, continuous-learning environment with smart, dedicated colleagues building the next generation of high-performance computing architecture

Similar Jobs

More Jobs at Netpreme

More Technical Services Jobs

Find similar AI Infrastructure Engineer jobs: