Full Job Description
We are looking for a Senior Staff Linux Runtime Software Engineer to own the user-mode components that interface with the Fabric and support the effcc compiler. This includes owning the user-mode driver, interfacing with our kernel mode driver, as well as defining the HAL interface exposed by the Fabric ABI. You will be working within pre-silicon environments, involving QEMU, functional simulation, emulation, and FPGAs to ensure applications run on day one.
This is a key role for the System Software team. As such, you will work closely with the compiler, kernel driver, architecture, micro-architecture, and verification teams during all phases of development.
Key Responsibilities
Own the Runtime Library and User-Mode Driver
• Design, implement, and maintain the user-mode runtime driver for the Fabric accelerator. This is a critical shared library which sits between the compiler and the Fabric ABI. This library implements context management, program loading, resource management, command queue construction and submission, descriptor allocation and setup, streaming, completion queue/event synchronization, error handling, logging, memory allocation, as well as general library support.
• Design and implement all user-mode interfaces between the compiler, the Fabric ABI, and the kernel mode driver.
• Define the in-memory representations generated by the compiler and consumed by the Fabric.
• Design and implement the underlying runtime library support and interfaces needed by 3rd party ML framework backends used by PyTorch, ONNX, and/or TensorFlow.
• Lead the user-mode efforts to bring up on first silicon. This includes being comfortable with functional simulation, emulation and FPGA environments.
• Establish runtime test coverage, including API unit and conformance tests, emulation-based regression suites, performance benchmarks, and hardware-in-the-loop validation.
Required Qualifications
• B.S. in Electrical/Computer Engineering, Computer Science, or equivalent; M.S. a plus.
• 7+ years of systems software experience, including significant hands-on Linux user-mode driver as well as language runtime library development for accelerators (GPUs, NPUs, DSPs, FPGAs, or similar).
• Experience building or contributing to a production accelerator runtime software stack, from the driver API up through the user-facing runtime library.
• Solid understanding of accelerator memory models, in particular unified memory systems, allocators, and zero-copy data paths.
• Experience designing public APIs exposed by runtime libraries: ABI stability, versioning, error models, thread safety, and documentation.
• Experience working with compiler teams on binary formats, loaders, launch ABIs, and kernel metadata, for ahead-of-time or just-in-time compilation.
• Experience integrating runtimes with language runtimes or ML frameworks: Python bindings, framework backends or execution providers, and C/C++ application interfaces.
• Hands-on experience with QEMU, emulation, or simulation environments during pre-silicon development.
• Experience designing interfaces between kernel and user space: ABI and HAL design, synchronization, memory sharing, and error handling.
• Strong user-mode debugging skills across crash analysis, sanitizers, tracing, profiling, and race and synchronization diagnosis.
• Working knowledge of ARM/AArch64 Linux systems and the Linux driver model.
• Expert C and C++ and strong Python skills; comfortable working with AI-assisted development tools for code, analysis, and debugging.
Desired Qualifications
• Experience contributing to CUDA Runtime (cudart), ROCm (HIP, ROCr, HSA), the Tenstorrent runtime (TT-Metalium), oneAPI Level Zero, OpenCL, or another accelerator runtime software stack.
• Familiarity with LLVM or MLIR, and with how compilers for accelerators hand off to runtimes.
• Knowledge of dataflow, spatial, or other novel compute architectures, and of how compilers and runtimes target them.
• Experience on a first-generation silicon product, where the software stack and the hardware matured together.
• Experience with FPGA prototyping and hardware emulation platforms (Palladium, Zebu, Veloce, or similar).
We offer a competitive base salary for this role, generally ranging from $220,000 to $270,000, plus a 10% annual bonus, meaningful equity, and comprehensive benefits. Final compensation will depend on experience, level, and location, with flexibility for the right candidate.