Software Engineer, Inference Runtime

LM Studio

• $135K — $160K *
Consumer Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Significant experience building production ML systems or performance-sensitive infrastructure
  • Strong programming skills in Python and C++
  • Deep understanding of transformer architectures and model inference mechanics
  • Experience profiling CPU/GPU workloads, focusing on compute, memory, and data management
  • Familiarity with PyTorch and inference systems like llama.cpp and TensorRT-LLM
  • Excellent debugging skills across model code and runtime internals
  • Strong sense of personal accountability for work quality and performance

Responsibilities

  • Enhance and maintain inference stack both on-device and in the cloud
  • Implement new model architectures and support multimodal models
  • Optimize latency, throughput, and resource utilization across various runtimes
  • Develop runtime capabilities for loading, batching, and scheduling models
  • Conduct benchmarking and troubleshooting of inference stack issues
  • Contribute to open-source projects like llama.cpp and MLX

Benefits

  • Great medical, vision, and dental plans
  • Catered lunches and dinners provided in-office
  • Flexible paid time off (PTO)
  • Work from home options available
  • Bright and inviting office in SoHo, NYC
Full Job Description
The Role

We are looking for an Inference Runtime Software Engineer to push forward LM Studio's inference stack on-device and in the cloud. You will integrate new inference engines and runtime capabilities, bring up new open-weight models and modalities, and optimize model execution for a wide range of CPU and GPU targets. You will also contribute improvements to the open-source projects we build on.

Qualifications
  • Significant experience building production ML systems, inference runtimes, or performance-sensitive infrastructure
  • Strong programming ability in Python and C++
  • Deep understanding of transformer architectures and the mechanics of model inference
  • Experience profiling CPU or GPU workloads and reasoning about compute, memory, synchronization, and data movement
  • Experience with PyTorch and inference systems such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM
  • Strong debugging instincts across model code, runtime internals, operating systems, and CPU or GPU execution
  • Takes personal responsibility for the correctness and performance of their work

Bonus Qualifications
  • Past contributions to open-source inference runtime projects such as llama.cpp, MLX, ExecuTorch, vLLM, SGLang, or TensorRT-LLM

Responsibilities
  • Maintain and push forward our inference stack on-device and in the cloud
  • Bring up new model architectures and multimodal models
  • Improve latency, throughput, memory use, and reliability across CPU, CUDA, Metal, Vulkan, and ROCm runtimes
  • Build runtime capabilities for model loading, batching, scheduling, caching, and distributed execution
  • Benchmark and diagnose correctness and performance problems across the inference stack
  • Contribute upstream to open-source projects such as llama.cpp and MLX

Benefits
  • Competitive salary and equity grants
  • Great medical, vision, dental healthcare plans
  • Catered team lunch / expensed dinners in the office
  • Flexible PTO
  • Flexible WFH
  • Sun-drenched office in SoHo in NYC

Similar Jobs

More Jobs at LM Studio

More Consumer Technology Jobs

Find similar Software Engineer, Inference Runtime jobs: