ML Systems Engineer

SEMI$150K — $350K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Experience with large-scale ML systems and GPU computing optimization.
  • Strong proficiency in Python and C++/CUDA; experience with SGLang, vLLM, PyTorch, or similar frameworks.
  • Deep understanding of GPU architecture and parallel computing paradigms.
  • Experience deploying LLMs in production with model serving and distributed inference.
  • Strong systems-level debugging and profiling skills across various stack layers.
  • Familiarity with distributed computing frameworks (Ray) is a plus.

Responsibilities

  • Design and optimize LLM inference systems across multi-node clusters to enhance throughput and reduce latency.
  • Implement and benchmark inference optimizations for production workloads.
  • Profile and analyze inference bottlenecks, considering GPU kernel execution and memory constraints.
  • Build evaluation harnesses and frameworks to measure various performance metrics.
  • Collaborate with research scientists on integrating model architectures into production systems.
  • Investigate and apply techniques from research and open-source to improve inference performance.

Benefits

  • Work on cutting-edge LLM optimization with real-world impact.
  • Access to substantial GPU resources for experimentation.
  • Collaborate with a world-class AI research and engineering team.
  • Shape AI system performance for leading semiconductor companies.
  • Unlimited PTO and comprehensive benefits (medical, vision, dental, 401k).
  • Two engineering-centered offices with amenities like free parking and meals.
Full Job Description
Position Overview

We are seeking an ML Systems Engineer to optimize the performance and efficiency of large language model inference powering our agentic AI platform. This is a technical role focused on low-level systems optimization. You will implement performance optimizations, build evaluation harnesses, and architect multi-node clusters for training and inference that push the limits of LLM throughput and latency. Your work will directly impact the responsiveness and cost-efficiency of AI agents used by leading semiconductor companies to design chips.

Key Responsibilities
  • Design, deploy, and optimize LLM inference systems across multi-node clusters, maximizing throughput and minimizing latency for production workloads.
  • Implement and benchmark concrete inference optimizations.
  • Profile and analyze inference bottlenecks at the systems level-from GPU kernel execution to memory bandwidth constraints.
  • Build robust evaluation harnesses and benchmarking frameworks that measure accuracy, throughput, latency, and resource utilization across various parallelism strategies.
  • Collaborate with research scientists to integrate new model architectures and optimizations into production inference infrastructure.
  • Investigate and apply emerging techniques from research papers and open-source projects to continuously improve inference performance.
Qualifications
  • B.S., M.S., or PhD in Computer Science, Electrical Engineering, or related field (or equivalent experience).
  • Experience with large-scale ML systems, GPU computing, or high-performance inference optimization.
  • Strong proficiency in Python and C++/CUDA; hands-on experience with SGLang, vLLM, PyTorch, or similar inference frameworks.
  • Deep understanding of GPU architecture, memory hierarchies, and parallel computing paradigms.
  • Experience deploying and optimizing LLMs in production: model serving, batching strategies, distributed inference, or quantization.
  • Strong systems-level debugging and profiling skills; comfort working at multiple layers of the stack from CUDA kernels to application logic.
  • Familiarity with distributed computing frameworks (Ray, multi-node training/inference) is a plus.
  • Self-directed problem solver who is interested in working on ambitious optimization challenges.
Why Join Us
  • Work on cutting-edge LLM inference optimization problems with real-world production impact.
  • Access to substantial GPU compute resources for experimentation and benchmarking.
  • Collaborate with a world-class team spanning AI research, systems engineering, and EDA.
  • Shape the performance characteristics of AI systems used by leading semiconductor companies.
What we offer
  • $150K/yr - $350K/yr + Offers Equity. We are open to discuss above-scale compensation with exceptional candidates on a case-by-case basis.
  • Unlimited PTO and full benefits (medical, vision, dental, 401k).
  • Two engineering-centric offices with free parking, private gym, and free lunch, drinks and snacks.

About SEMI

SEMI is a global industry association that represents companies involved in the manufacturing and supply chain of the electronics industry. The association was founded in 1970 and is headquartered in Milpitas, California. SEMI provides its members with networking opportunities, industry research, and advocacy on issues affecting the industry. The association also hosts industry events and conferences, including SEMICON West and SEMICON China. SEMI has offices and operations in North America, Europe, Japan, Korea, Taiwan, China, and Southeast Asia.
Learn more about SEMI
Size
200 employees
Industry

Similar Jobs

More Jobs at SEMI

More Information Technology Jobs

Find similar ML Systems Engineer jobs: