Mount Thor is hiring a founding inference engineer to build a production inference platform on Apple Silicon. You will own the stack from GPU kernels and model execution to distributed scheduling, networking, and customer-facing serving systems.
The RoleYou will establish the technical foundation for inference at Mount Thor. You will define the architecture, write the performance-critical code, and take the system through deployment and production operation.
Your scope spans the complete path from an incoming inference request to execution on the hardware and a response back to the customer. You will work across kernels, compilers, inference engines, memory management, model parallelism, networking, scheduling, and serving APIs. You will make decisions about where to extend existing open-source systems and where to build new components.
You will work closely with Platform, Fleet, and Network Infrastructure engineering to turn individual machines into a reliable inference service. You will own the tradeoffs among latency, throughput, model quality, reliability, and cost, and use measurements to guide the work.
In this role you will- Build and optimize the inference engine. Implement model execution, continuous batching, prefill and decode scheduling, KV-cache management, prefix caching, and memory allocation. Bring new model architectures into production and implement optimizations such as quantization, speculative decoding, and chunked prefill.
- Write performance-critical kernels and runtime code. Implement and optimize operations such as matrix multiplication, attention, and mixture-of-experts execution. Work across Metal, MLX, and compiler and runtime internals to improve memory access, operator fusion, synchronization, and CPU/GPU execution. Validate numerical correctness alongside performance.
- Design for Apple Silicon. Build around unified memory, memory bandwidth, compute capabilities, and operating-system behavior. Profile execution on real hardware, diagnose bottlenecks, and choose model representations and execution strategies based on measured results.
- Build distributed inference and its communication layer. Implement model sharding and parallel execution across machines. Develop and optimize collective communication, data movement, and the overlap between computation and communication. Work with network engineers on transport performance, topology, and failure handling.
- Own inference scheduling and routing. Build request queues, admission control, cache-aware routing, model placement, and autoscaling. Manage capacity and tenant fairness while meeting latency and throughput targets across changing traffic patterns and model sizes.
- Ship a production inference service. Build model loading and deployment workflows, serving APIs, streaming responses, cancellation, and backpressure. Integrate the runtime with fleet and platform systems. Own observability, safe releases, incident diagnosis, and recovery for the inference stack.
- Make performance reproducible. Build benchmarks and profiling tools that measure time to first token, inter-token latency, tail latency, throughput, memory use, power, and cost per token. Test realistic workloads, concurrency, and context lengths. Catch performance and model-quality regressions before release.
- Set the engineering direction. Turn customer workloads into technical priorities, choose and contribute to open-source projects, and establish clear interfaces across the stack. Use coding agents to investigate, implement, and validate changes, backed by correctness checks and reproducible benchmarks.
What you bring- A record of building and optimizing production inference engines, GPU compute software, or distributed ML systems, with substantial depth in at least one and hands-on work across multiple layers.
- Strong systems programming skills in C++ or Rust, proficiency in Python, and experience working inside performance-critical libraries and runtimes.
- A practical understanding of transformer inference, including attention, prefill and decode, batching, KV caches, quantization, and their compute and memory costs.
- Experience with GPU programming and performance analysis, including memory hierarchies, parallel execution, synchronization, and numerical precision.
- Strong distributed-systems fundamentals, including scheduling, concurrency, networking, failure handling, and resource management.
- Experience taking software from architecture through production deployment, debugging, and ongoing operation.
- The ability to turn an ambiguous performance problem into a measured bottleneck, an implementation, and a verified improvement.
- The judgment and ownership to establish a new technical area, prioritize the work, and explain tradeoffs clearly to engineers and customers.
Bonus Skills- Built with or contributed to MLX, MLX-LM, llama.cpp, vLLM, SGLang, TensorRT-LLM, or similar systems.
- Written Metal kernels or worked on Apple GPU performance, macOS internals, unified-memory management, or compiler infrastructure.
- Implemented tensor, pipeline, or expert parallelism; distributed KV caches; or disaggregated prefill and decode.
- Built collective communication libraries or optimized inference over Ethernet, Thunderbolt, or RDMA transports.
- Shipped a multi-tenant inference service, model-serving control plane, or serverless compute product.
- Helped establish an engineering team or core infrastructure product at an early-stage company.