What You'll Do Here- Build the host-side interface library - device memory management, DMA, streams and events, sync primitives - that every compiler-emitted program runs on top of
- Own and extend the executable format: the compiler→runtime contract, its versioning, the weight and quantization layouts that let compiler and runtime evolve independently
- Design the custom-kernel ABI - calling convention, sync semantics, lifecycle - and the host-side marshaling layer (DLPack, the buffer protocol, numpy) that gets Python tensors to the device
- Build Python bindings via PyO3, with a C-ABI shim as the alternative integration path for downstream consumers
- Build the LLM inference serving stack - paged KV cache, continuous batching, request scheduling, token streaming - and the cluster orchestration primitives underneath it
- Bring up interconnect topology from the host and own the failure-detection and clean-teardown path for stop-restructure-resume recovery across racks
- Design what the chip exposes to host-side profilers and debuggers - perf counters, traces, and the Python surfaces ML engineers actually use - and hit measurable performance targets on runtime overhead and serving throughput
Who You Are- Strong experience in a systems programming language - Rust, C, C++, or Go - including memory management, allocator design, and FFI/ABI work
- Have built Python interop layers in production (PyO3, ctypes, pybind11, or equivalent C-ABI bridging)
- Have designed and maintained API or ABI contracts between teams - versioning, evolution, breaking-change discipline - not just consumed someone else's
- Hands-on with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or comparable) - enough to reason about device memory, async execution, and kernel launch
- ML-systems literate - comfortable with the training and inference loop, what collectives do, what a tensor layout is. Research depth not required.
Bonus Points If You Have- LLM inference internals - vLLM, TensorRT-LLM, or SGLang (paged attention, scheduler design)
- Rust at depth, including proc macros, unsafe with soundness reasoning, and complex lifetime/trait work
- Custom allocator design (slab, paged, arena) or other low-level memory work
- ML framework integration experience (PyTorch custom backends, JAX/XLA, ONNX runtime)
- Profiler or tracing infrastructure work (perfetto, Nsight, or a custom stack)
- Driver-adjacent or kernel-bypass work, or prior new-silicon bring-up
CompensationThe US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job-related skills, and relevant education and training. Career length is only a guideline for compensation.
- Early Career - $160,000 - $250,000 + equity
- Mid Career - $175,000 - $362,500 + equity
- Senior Career - $250,000 - $475,000 + equity
What We Offer- Time off: 4 weeks PTO (accrued) + 12 company Holidays + up to 3 weeks remote work
- Health: Company-subsidized Medical (Kaiser or Anthem) for employees & dependents, Guardian Dental and Vision insurances for employee & dependents, and life insurance (employee only), plus HSA and FSA offerings via Lively.
- Financial Wellbeing: Choose from Roth IRA/ 401K (or both) retirement plans with up to 5% company contribution to 401K (even if you don't contribute). Also, 100% company-paid life insurance (up to $300K) and long-term disability insurances.
- Professional Development: $1500 Professional Development Budget (per year)
- Team Meals: MatX provides onsite team lunch & dinner Monday - Friday, with your choice of ordering via WeBox, Specialty's or via our reimbursement system
- Commute on Us: Commute on our company Uber account, or reimburse your train rides. Either way, we pay 100% for your daily commute.
- MatX E[x]tras: $50/mo to use on the perk you value most
- Cell & Internet Reimbursement: $35/mo for cellular and $40/mo for wifi
- Mental Wellbeing: 100% paid mental health benefit via SpringHealth and Guardian EAP.
- Support to Parents: Up to 12 weeks paid parental leave regardless of path to parenthood, 10 weeks pregnancy disability leave, flexible return-to-work hours, and Benepass reproductive health & parental benefit.
- AI Resources: Up to $20K/month plus a dedicated internal AI Tooling Team to support your productivity
All candidates must be authorized to work in the United States and work from our offices in Mountain View Tuesdays-Thursdays.