Member of Technical Staff, Inference

Mount Thor

$150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Experience in building and optimizing production inference engines or distributed ML systems.
  • Strong systems programming skills in C++ or Rust, and proficiency in Python.
  • Solid understanding of transformer inference and its compute and memory costs.
  • Knowledge of GPU programming and performance analysis techniques.
  • Strong foundations in distributed systems, including concurrency and failure handling.
  • Experience from architecture through to deployment and ongoing operation of software.
  • Ability to analyze and resolve ambiguous performance issues.

Responsibilities

  • Build and optimize the inference engine with model execution and memory management.
  • Write performance-critical kernels and runtime code for operations like matrix multiplication.
  • Design systems specifically for Apple Silicon, profiling hardware performance and diagnosing issues.
  • Implement distributed inference and optimize communication across machines.
  • Own the inference scheduling, routing, and manage capacity and latency targets.
  • Ship a production-grade inference service integrating various workflows and systems.
  • Develop benchmarks and profiling tools to ensure performance reliability and traceability.
  • Set and guide the engineering direction based on customer workloads and performance metrics.

Benefits

  • Opportunity to influence the engineering direction from the ground up.
  • Work in an innovative environment focused on Apple Silicon and production inference systems.
  • Collaborate closely with multiple engineering teams to develop complex systems.
  • Be part of a dynamic team working on cutting-edge technology.
Full Job Description
Mount Thor is hiring a founding inference engineer to build a production inference platform on Apple Silicon. You will own the stack from GPU kernels and model execution to distributed scheduling, networking, and customer-facing serving systems.
The Role

You will establish the technical foundation for inference at Mount Thor. You will define the architecture, write the performance-critical code, and take the system through deployment and production operation.

Your scope spans the complete path from an incoming inference request to execution on the hardware and a response back to the customer. You will work across kernels, compilers, inference engines, memory management, model parallelism, networking, scheduling, and serving APIs. You will make decisions about where to extend existing open-source systems and where to build new components.

You will work closely with Platform, Fleet, and Network Infrastructure engineering to turn individual machines into a reliable inference service. You will own the tradeoffs among latency, throughput, model quality, reliability, and cost, and use measurements to guide the work.
In this role you will
  • Build and optimize the inference engine. Implement model execution, continuous batching, prefill and decode scheduling, KV-cache management, prefix caching, and memory allocation. Bring new model architectures into production and implement optimizations such as quantization, speculative decoding, and chunked prefill.
  • Write performance-critical kernels and runtime code. Implement and optimize operations such as matrix multiplication, attention, and mixture-of-experts execution. Work across Metal, MLX, and compiler and runtime internals to improve memory access, operator fusion, synchronization, and CPU/GPU execution. Validate numerical correctness alongside performance.
  • Design for Apple Silicon. Build around unified memory, memory bandwidth, compute capabilities, and operating-system behavior. Profile execution on real hardware, diagnose bottlenecks, and choose model representations and execution strategies based on measured results.
  • Build distributed inference and its communication layer. Implement model sharding and parallel execution across machines. Develop and optimize collective communication, data movement, and the overlap between computation and communication. Work with network engineers on transport performance, topology, and failure handling.
  • Own inference scheduling and routing. Build request queues, admission control, cache-aware routing, model placement, and autoscaling. Manage capacity and tenant fairness while meeting latency and throughput targets across changing traffic patterns and model sizes.
  • Ship a production inference service. Build model loading and deployment workflows, serving APIs, streaming responses, cancellation, and backpressure. Integrate the runtime with fleet and platform systems. Own observability, safe releases, incident diagnosis, and recovery for the inference stack.
  • Make performance reproducible. Build benchmarks and profiling tools that measure time to first token, inter-token latency, tail latency, throughput, memory use, power, and cost per token. Test realistic workloads, concurrency, and context lengths. Catch performance and model-quality regressions before release.
  • Set the engineering direction. Turn customer workloads into technical priorities, choose and contribute to open-source projects, and establish clear interfaces across the stack. Use coding agents to investigate, implement, and validate changes, backed by correctness checks and reproducible benchmarks.
What you bring
  • A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
  • Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
  • A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
  • Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
  • Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
  • Experience taking software from architecture through production deployment, debugging, and ongoing operation.
  • The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
  • The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.
Bonus Skills
  • Built with or contributed to MLX, MLX-LM, llama.cpp, vLLM, SGLang, TensorRT-LLM, or similar systems.
  • Written Metal kernels or worked on Apple GPU performance, macOS internals, unified-memory management, or compiler infrastructure.
  • Implemented tensor, pipeline, or expert parallelism; distributed KV caches; or disaggregated prefill and decode.
  • Built collective communication libraries or optimized inference over Ethernet, Thunderbolt, or RDMA transports.
  • Shipped a multi-tenant inference service, model-serving control plane, or serverless compute product.
  • Helped establish an engineering team or core infrastructure product at an early-stage company.

Similar Jobs

More Jobs at Mount Thor

More Enterprise Technology Jobs

Find similar Member of Technical Staff, Inference jobs: