Member of Technical Staff, Inference

Mount Thor

$150K — $180K *
US-AnywhereRemote in San Francisco, CA
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in building or optimizing production inference engines or distributed ML systems.
  • Strong systems programming skills in C++ or Rust, with proficiency in Python.
  • Deep understanding of transformer inference mechanisms, including attention and batching.
  • Expertise in GPU programming with a focus on performance analysis and optimization.
  • Thorough knowledge of distributed systems principles like scheduling, concurrency, and resource management.
  • Experience in the full software lifecycle from architecture to production deployment and operation.
  • Ability to diagnose performance issues and implement measurable improvements.

Responsibilities

  • Build and optimize the inference engine and its performance-critical components.
  • Write and optimize runtime code and kernels for operations like matrix multiplication and attention.
  • Design systems specifically for Apple Silicon, focusing on memory and execution optimizations.
  • Implement distributed inference solutions and enhance communication layers for parallel execution.
  • Manage inference scheduling, routing, and autoscaling while ensuring latency and throughput targets.
  • Ship robust production inference services with seamless deployment workflows and observability features.
  • Establish benchmarks and tools to ensure reproducible performance and catch regressions.

Benefits

  • Opportunity to work on cutting-edge AI and ML technologies.
  • Collaboration with a talented team at a pioneering AI startup.
  • Chance to define and influence the technical direction of a new product.
  • Work in a dynamic environment fostering innovation and strong engineering practices.
  • Access to career growth opportunities in a fast-paced industry.
Full Job Description
Mount Thor is hiring a founding inference engineer to build a production inference platform on Apple Silicon. You will own the stack from GPU kernels and model execution to distributed scheduling, networking, and customer-facing serving systems.
The Role

You will establish the technical foundation for inference at Mount Thor. You will define the architecture, write the performance-critical code, and take the system through deployment and production operation.

Your scope spans the complete path from an incoming inference request to execution on the hardware and a response back to the customer. You will work across kernels, compilers, inference engines, memory management, model parallelism, networking, scheduling, and serving APIs. You will make decisions about where to extend existing open-source systems and where to build new components.

You will work closely with Platform, Fleet, and Network Infrastructure engineering to turn individual machines into a reliable inference service. You will own the tradeoffs among latency, throughput, model quality, reliability, and cost, and use measurements to guide the work.
In this role you will
  • Build and optimize the inference engine. Implement model execution, continuous batching, prefill and decode scheduling, KV-cache management, prefix caching, and memory allocation. Bring new model architectures into production and implement optimizations such as quantization, speculative decoding, and chunked prefill.
  • Write performance-critical kernels and runtime code. Implement and optimize operations such as matrix multiplication, attention, and mixture-of-experts execution. Work across Metal, MLX, and compiler and runtime internals to improve memory access, operator fusion, synchronization, and CPU/GPU execution. Validate numerical correctness alongside performance.
  • Design for Apple Silicon. Build around unified memory, memory bandwidth, compute capabilities, and operating-system behavior. Profile execution on real hardware, diagnose bottlenecks, and choose model representations and execution strategies based on measured results.
  • Build distributed inference and its communication layer. Implement model sharding and parallel execution across machines. Develop and optimize collective communication, data movement, and the overlap between computation and communication. Work with network engineers on transport performance, topology, and failure handling.
  • Own inference scheduling and routing. Build request queues, admission control, cache-aware routing, model placement, and autoscaling. Manage capacity and tenant fairness while meeting latency and throughput targets across changing traffic patterns and model sizes.
  • Ship a production inference service. Build model loading and deployment workflows, serving APIs, streaming responses, cancellation, and backpressure. Integrate the runtime with fleet and platform systems. Own observability, safe releases, incident diagnosis, and recovery for the inference stack.
  • Make performance reproducible. Build benchmarks and profiling tools that measure time to first token, inter-token latency, tail latency, throughput, memory use, power, and cost per token. Test realistic workloads, concurrency, and context lengths. Catch performance and model-quality regressions before release.
  • Set the engineering direction. Turn customer workloads into technical priorities, choose and contribute to open-source projects, and establish clear interfaces across the stack. Use coding agents to investigate, implement, and validate changes, backed by correctness checks and reproducible benchmarks.
What you bring
  • A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
  • Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
  • A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
  • Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
  • Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
  • Experience taking software from architecture through production deployment, debugging, and ongoing operation.
  • The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
  • The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.
Bonus Skills
  • Built with or contributed to MLX, MLX-LM, llama.cpp, vLLM, SGLang, TensorRT-LLM, or similar systems.
  • Written Metal kernels or worked on Apple GPU performance, macOS internals, unified-memory management, or compiler infrastructure.
  • Implemented tensor, pipeline, or expert parallelism; distributed KV caches; or disaggregated prefill and decode.
  • Built collective communication libraries or optimized inference over Ethernet, Thunderbolt, or RDMA transports.
  • Shipped a multi-tenant inference service, model-serving control plane, or serverless compute product.
  • Helped establish an engineering team or core infrastructure product at an early-stage company.

Similar Jobs

More Jobs at Mount Thor

More Enterprise Technology Jobs

Find similar Member of Technical Staff, Inference jobs: