Member of Technical Staff, ML Inference Engineering

Sanas.ai

• $160K — $190K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience writing high-quality, high-performance code.
  • Familiarity with NVIDIA GPU architecture and CUDA.
  • Fluency in the LLM serving stack, covering kernels to autoscaling.
  • Research or systems background in LLM, speech inference with demonstrable projects.
  • A history of developing systems or research for others to utilize.

Responsibilities

  • Optimize system and GPU performance for AI workloads across multi-node setups.
  • Analyze and enhance latency, throughput, memory usage, and compute efficiency.
  • Profile system performance to identify and resolve bottlenecks.
  • Implement low-level optimizations using CUDA and performance tools.
  • Establish performance benchmarking and monitoring infrastructure.
  • Own and advance the inference engine for scale and reliability.
  • Develop runtime inference services for large-scale AI applications.

Benefits

  • Work closely with cutting-edge technology in a challenging environment.
  • Opportunity to shape core infrastructure and architect solutions.
  • Be part of a team that promotes peer growth and learning.
  • Engage in hands-on engineering and optimization activities.
Full Job Description
About the Role

Sanas is bringing real-time speech and language models on-premise - deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Inference Optimizations
• Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.
• Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.
• Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.
• Drive model-level execution improvements, including mixed precision and advanced quantization strategies.

Serving Optimizations
• Write and optimize custom inference server backends to handle complex business logic, scripting, and state management.
• Own and scale our overarching model serving infrastructure across multi-GPU, multi-node deployments to meet strict on-premise latency budgets.
• Build robust, fault-tolerant runtime services, complete with comprehensive performance benchmarking and monitoring.

Requirements
• 8+ years of experience writing high-quality, high-performance code, including at least 5 years focused on machine learning systems.
• Experience with one of C/C++/Rust and Python.
• Deep familiarity with modern NVIDIA GPU architectures (e.g., Ada Lovelace, Blackwell), CUDA, and low-level system profiling.
• Hands-on experience with ML compilers and optimization frameworks (e.g., Apache TVM, TorchInductor/Dynamo, TensorRT, Triton).
• Experience building or extending model serving infrastructure, specifically writing custom C++ backends for Triton Inference Server (or similar serving engines).
• Fluency in the AI serving stack, from kernels and quantization up to schedulers, state management, and autoscaling.
• A record of shipping research or systems that other people build on, whether in a lab or in industry.

Nice-to-have:
• Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes.
• A research-leaning or systems background in Speech (STT, TTS, S2S) or LLM inference, with work you can point to.
• Experience operating large-scale, on-premise AI training or inference clusters, including bare-metal provisioning and Kubernetes management.
• Familiarity with high-performance networking (InfiniBand or RoCE), distributed storage systems, and hardware health monitoring.
• Experience maintaining or contributing to open-source ML or systems projects.

Similar Jobs

More Jobs at Sanas.ai

More Information Technology Jobs

Find similar Member of Technical Staff, ML Inference Engineering jobs: