Runtime Engineer

MatX

$160K — $475K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong experience in systems programming languages like Rust, C, C++, or Go
  • Proven experience building Python interop layers with tools like PyO3 or ctypes
  • Experience designing and maintaining API or ABI contracts
  • Hands-on with accelerator programming models like CUDA or oneAPI
  • Familiar with machine learning systems concepts, particularly in training and inference.

Responsibilities

  • Build host-side interface libraries for compiler environments
  • Extend executable formats for independent compiler and runtime evolution
  • Design custom-kernel ABI and marshaling layers for data transfer
  • Create Python bindings and alternative C-ABI integration paths
  • Develop LLM inference serving stack including caching and batching
  • Manage interconnect topology and recovery mechanisms across racks
  • Design host-side profiling and debugging tools with performance targets.

Benefits

  • 4 weeks PTO plus 12 holidays and up to 3 weeks remote work
  • Company-subsidized health insurance for employees and dependents
  • Investment in financial wellbeing with 401K contributions
  • Annual professional development budget of $1500
  • Onsite team meals provided during workdays
  • Fully covered transportation costs for commutes
  • Monthly perks allowance and reimbursements for internet and cell expenses
  • Comprehensive mental health benefits paid for by the company
  • Generous parental leave policies and flexible return-to-work options
  • Up to $20K/month in AI resources and support team for productivity.
Full Job Description
What You'll Do Here
  • Build the host-side interface library - device memory management, DMA, streams and events, sync primitives - that every compiler-emitted program runs on top of
  • Own and extend the executable format: the compiler→runtime contract, its versioning, the weight and quantization layouts that let compiler and runtime evolve independently
  • Design the custom-kernel ABI - calling convention, sync semantics, lifecycle - and the host-side marshaling layer (DLPack, the buffer protocol, numpy) that gets Python tensors to the device
  • Build Python bindings via PyO3, with a C-ABI shim as the alternative integration path for downstream consumers
  • Build the LLM inference serving stack - paged KV cache, continuous batching, request scheduling, token streaming - and the cluster orchestration primitives underneath it
  • Bring up interconnect topology from the host and own the failure-detection and clean-teardown path for stop-restructure-resume recovery across racks
  • Design what the chip exposes to host-side profilers and debuggers - perf counters, traces, and the Python surfaces ML engineers actually use - and hit measurable performance targets on runtime overhead and serving throughput
Who You Are
  • Strong experience in a systems programming language - Rust, C, C++, or Go - including memory management, allocator design, and FFI/ABI work
  • Have built Python interop layers in production (PyO3, ctypes, pybind11, or equivalent C-ABI bridging)
  • Have designed and maintained API or ABI contracts between teams - versioning, evolution, breaking-change discipline - not just consumed someone else's
  • Hands-on with at least one accelerator programming model (CUDA, ROCm, oneAPI Level Zero, TPU, or comparable) - enough to reason about device memory, async execution, and kernel launch
  • ML-systems literate - comfortable with the training and inference loop, what collectives do, what a tensor layout is. Research depth not required.
Bonus Points If You Have
  • LLM inference internals - vLLM, TensorRT-LLM, or SGLang (paged attention, scheduler design)
  • Rust at depth, including proc macros, unsafe with soundness reasoning, and complex lifetime/trait work
  • Custom allocator design (slab, paged, arena) or other low-level memory work
  • ML framework integration experience (PyTorch custom backends, JAX/XLA, ONNX runtime)
  • Profiler or tracing infrastructure work (perfetto, Nsight, or a custom stack)
  • Driver-adjacent or kernel-bypass work, or prior new-silicon bring-up
Compensation

The US base salary for this full-time position is determined based on a variety of factors including role, experience, location, job-related skills, and relevant education and training. Career length is only a guideline for compensation.
  • Early Career - $160,000 - $250,000 + equity
  • Mid Career - $175,000 - $362,500 + equity
  • Senior Career - $250,000 - $475,000 + equity


What We Offer
  • Time off: 4 weeks PTO (accrued) + 12 company Holidays + up to 3 weeks remote work
  • Health: Company-subsidized Medical (Kaiser or Anthem) for employees & dependents, Guardian Dental and Vision insurances for employee & dependents, and life insurance (employee only), plus HSA and FSA offerings via Lively.
  • Financial Wellbeing: Choose from Roth IRA/ 401K (or both) retirement plans with up to 5% company contribution to 401K (even if you don't contribute). Also, 100% company-paid life insurance (up to $300K) and long-term disability insurances.
  • Professional Development: $1500 Professional Development Budget (per year)
  • Team Meals: MatX provides onsite team lunch & dinner Monday - Friday, with your choice of ordering via WeBox, Specialty's or via our reimbursement system
  • Commute on Us: Commute on our company Uber account, or reimburse your train rides. Either way, we pay 100% for your daily commute.
  • MatX E[x]tras: $50/mo to use on the perk you value most
  • Cell & Internet Reimbursement: $35/mo for cellular and $40/mo for wifi
  • Mental Wellbeing: 100% paid mental health benefit via SpringHealth and Guardian EAP.
  • Support to Parents: Up to 12 weeks paid parental leave regardless of path to parenthood, 10 weeks pregnancy disability leave, flexible return-to-work hours, and Benepass reproductive health & parental benefit.
  • AI Resources: Up to $20K/month plus a dedicated internal AI Tooling Team to support your productivity

All candidates must be authorized to work in the United States and work from our offices in Mountain View Tuesdays-Thursdays.

Similar Jobs

More Information Technology Jobs

Find similar Runtime Engineer jobs: