Member of Technical Staff - Inference

Sail Research

• $150K — $180K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in GPU programming or inference engine development
  • Deep knowledge of large language models (LLMs) and their mechanics
  • Familiarity with cutting-edge research in Machine Learning Systems (MLSys)
  • Experience with advanced GPU programming frameworks like Triton or CUTLASS
  • Strong communication skills, particularly in technical discussions

Responsibilities

  • Modify state-of-the-art inference engines like vLLM and SGLang
  • Analyze GPU time allocation during operations, explaining kernel launches effectively
  • Design innovative parallelism strategies for advanced hardware architectures
  • Develop and optimize GPU kernels for specific computational tasks
  • Benchmark inference performance and investigate cache hit rate enhancements

Benefits

  • Free meals provided in the office, leveraging the local food scene
  • High-quality hardware setup including Studio Displays for optimal work performance
  • Supportive office culture with a focus on productivity and leisure
  • Access to a friendly office pet for a welcoming work environment
  • Focus on investments in tools that enhance work efficiency
Full Job Description
In this role, you'll own token processing down to the lowest layers of the stack. You'll do things like: develop a new request scheduling strategy, achieve better communication/computation overlap, investigate novel schemes for increasing cache hit rates, or identify a better way to benchmark inference performance.

What you'll do
  • Modify and extend state-of-the-art inference engines like vLLM and SGLang, and work on our own internal engine.
  • Understand every microsecond of GPU time spent during a forward pass. You'll be able to explain every kernel launch on an nsys profile.
  • Design and implement exotic parallelism schemes to work with "interesting" hardware topologies.
  • Write and debug GPU kernels to excel in specific regimes, such as cascade attention

What we're looking for
  • Strong understanding of core LLM mechanics, like KV cache, mixture-of-experts, prefill vs. decode phases.
  • Interest in MLSys research - great ideas like speculative decoding and sparse attention come from research, that we need to follow closely.
  • Familiarity with modern, tile-based GPU programming, e.g. Triton, CUTLASS, ThunderKittens, etc. Or an interest in learning these!
  • Great interpersonal and technical communication. Please don't use LLMs to write prose. We desk-reject slopful cover letters and resumes.


Interview process
  1. Meet the CTO, who will ask about your experience, and share as much technical detail about Sail as you want to hear. This is the first step because we respect your time.
  2. Share an online whiteboard with a team member and work through a technical problem. We spend a lot of time at whiteboards, building intuition about complex systems together. It's a great way for us to see how you communicate technically, and a even better way for you to see what working at Sail is like.
  3. Come in to Sail's SF office for an interview day. Meet the whole team, and work on a bunch of problems that closely simulates the work we do daily. We'll also ask you to give us a 20-30min 'chalk talk' about an interesting problem you've worked on before.
  4. Offer. Once the team decides we want to work with you, we make a strong offer quickly and will be quite persistent over email/text/calls :)


Life at Sail

We work out of a beautiful, sunny office in downtown San Francisco. All meals are on us (and actually great; SF is a food paradise!). Everyone gets a Studio Display (or two) at their desk. We are serious about investing in anything that saves us time or energy. There are six different ways to make coffee or tea in the office. A friendly (hypoallergenic) black cat named Coco visits occasionally.

Similar Jobs

More Jobs at Sail Research

More Enterprise Technology Jobs

Find similar Member of Technical Staff - Inference jobs: