Full Job Description
About the Role
Sanas is bringing real-time speech and language models on-premise - deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.
We're looking for a deeply hands-on, experienced engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.
What You'll Do
Inference Optimizations
• Implement custom kernels and low-level optimizations to push the absolute limits of GPU compute.
• Apply graph optimization, operator fusion, and hardware-specific code generation at the ML compiler level.
• Profile, analyze, and resolve deep system bottlenecks to radically improve latency, throughput, and memory efficiency.
• Drive model-level execution improvements, including mixed precision and advanced quantization strategies.
Serving Optimizations
• Write and optimize custom inference server backends to handle complex business logic, scripting, and state management.
• Own and scale our overarching model serving infrastructure across multi-GPU, multi-node deployments to meet strict on-premise latency budgets.
• Build robust, fault-tolerant runtime services, complete with comprehensive performance benchmarking and monitoring.
Requirements
• 8+ years of experience writing high-quality, high-performance code, including at least 5 years focused on machine learning systems.
• Experience with one of C/C++/Rust and Python.
• Deep familiarity with modern NVIDIA GPU architectures (e.g., Ada Lovelace, Blackwell), CUDA, and low-level system profiling.
• Hands-on experience with ML compilers and optimization frameworks (e.g., Apache TVM, TorchInductor/Dynamo, TensorRT, Triton).
• Experience building or extending model serving infrastructure, specifically writing custom C++ backends for Triton Inference Server (or similar serving engines).
• Fluency in the AI serving stack, from kernels and quantization up to schedulers, state management, and autoscaling.
• A record of shipping research or systems that other people build on, whether in a lab or in industry.
Nice-to-have:
• Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes.
• A research-leaning or systems background in Speech (STT, TTS, S2S) or LLM inference, with work you can point to.
• Experience operating large-scale, on-premise AI training or inference clusters, including bare-metal provisioning and Kubernetes management.
• Familiarity with high-performance networking (InfiniBand or RoCE), distributed storage systems, and hardware health monitoring.
• Experience maintaining or contributing to open-source ML or systems projects.