Senior ML Systems Engineer, Inference

RunPod, Inc.

• $150K — $220K *
US-AnywhereRemote in United States
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of system engineering experience
  • Hands-on experience with vLLM, SGLang, or equivalent serving engines
  • Strong Python software engineering skills in performance-critical environments
  • Deep understanding of LLM inference performance drivers
  • Familiarity with modern inference optimization techniques
  • Experience in benchmarking with GPU profiling tools
  • Ability to effectively communicate findings in writing

Responsibilities

  • Define metrics for measuring inference performance and develop rigorous tooling for assessments
  • Diagnose and profile performance issues across the entire serving stack
  • Enhance serving efficiency for large models on both single and multi-node GPU setups
  • Develop production-ready runtimes and configurations from learned insights
  • Collaborate with product and infrastructure teams to define inference offerings
  • Stay updated on the inference ecosystem and contribute to open-source communities
  • Trace and resolve bottlenecks in serving engines beyond configuration tuning

Benefits

  • Competitive base pay with potential for significant salary range
  • Equity opportunities for all employees with stock options
  • Comprehensive medical, dental, and vision coverage
  • Flexible PTO to promote work-life balance
  • Predominantly remote working environment with a focus on inclusivity
  • $1,200 stipend for home office setup to support a productive workspace
  • Opportunity to join a dynamic team in the rapidly evolving AI infrastructure space
Full Job Description
We\'re looking for a ML Systems Engineer, Inference. We want Runpod to be the best place in the world to run LLM inference, meaning the fastest and the most cost-efficient. You\'ll lead that effort. You\'ll own LLM serving performance end to end. That means measuring it, understanding it, and improving it across models, hardware generations, and workloads. The work you ship will show up directly in the latency and cost our customers experience. This is a hands-on engineering role for someone who likes finding the real bottleneck and fixing it, then turning that fix into something that runs reliably in production.

Responsibilities
  • Define how we measure inference performance, including throughput, time to first token, inter-token latency, and cost per token, and build the tooling that makes those measurements rigorous and repeatable.
  • Profile and diagnose performance problems across the serving stack, from scheduling and memory management down to kernels and interconnect.
  • Improve serving efficiency for large, state-of-the-art models on single-node and multi-node GPU deployments.
  • Turn what you learn into production-ready runtimes, configurations, and defaults that customers benefit from automatically.
  • Work closely with product and infrastructure teams to shape how inference is offered on Runpod.
  • Keep up with the fast-moving inference ecosystem, including the open-source community, and decide what\'s worth adopting, what\'s worth building, and what\'s worth contributing back.
  • Trace bottlenecks in the serving engine/runtime and implement fixes when configuration tuning is not enough.


Requirements
  • 5+ years of professional system engineering experience.
  • Deep, hands-on experience with vLLM, SGLang (or a comparable serving engine) in production or at serious benchmark scale.
  • Strong software engineering skills in Python. You\'re comfortable working in large, performance-critical codebases.
  • A solid understanding of what drives LLM inference performance: batching, memory, parallelism, and the trade-offs between latency and throughput.
  • Experience with modern inference optimization techniques such as quantization, speculative decoding, or distributed serving.
  • Rigor in benchmarking and performance analysis, plus comfort with GPU profiling tools.
  • The ability to explain your results clearly in writing and turn them into decisions.


Preferred
  • Experience writing or tuning GPU kernels in CUDA or Triton.
  • Contributions to inference or ML systems projects.
  • Experience with multi-node GPU systems and high-speed networking.
  • Experience at a company where inference cost and latency were core business metrics.


What You\'ll Receive:
  • The competitive base pay for this position ranges from ($150,000 - $220,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate\'s experience, qualifications, and location
  • Meaningful equity in a fast-growing company- everyone on the team receives stock options - your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans
  • Flexible PTO- take the time you need to recharge
  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication
  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.
  • $1,200 Home Office & Equipment Stipend-We set you up for success from day one with gear and support to create your ideal workspace

Similar Jobs

More Jobs at RunPod, Inc.

More Consumer Technology Jobs

Find similar Senior ML Systems Engineer, Inference jobs: