4+ years of experience with compilers or runtime systems
Deep understanding of asynchronous and concurrent programming
4+ years of experience in C/C++ (C++14 or newer)
Familiarity with hardware architecture, including memory hierarchies
Knowledge of operating system kernel or hypervisor development
Responsibilities
Design and improve the multi-target runtime
Automate kernel generation using parallelization techniques
Prototype and explore innovative runtime ideas
Benchmark compiler outputs on various hardware
Analyze performance bottlenecks and build analysis tools
Collaborate with product team to enhance runtime architecture
Benefits
Opportunities to impact next-gen AI infrastructure
Work across the full tech stack from hardware to execution
Join a creative team focused on innovative infrastructure solutions
Includes competitive compensation with equity and bonuses
Comprehensive medical, dental, and vision coverage
Retirement savings plan and wellness benefits
Full Job Description
About the Role
We're looking for a Runtime Engineer to design and build the multi-target runtime that sits at the heart of our AI compiler stack. This is a systems-level role where you'll take the output of our optimizing compiler and make it execute - efficiently, correctly, and at scale - across a diverse landscape of hardware targets.
You'll work on low-level parallelization, kernel scheduling, and performance analysis, and collaborate closely with our compiler and product teams to push the boundaries of what's possible on modern AI hardware. What You'll Do
Design, develop, maintain, and improve our multi-target runtime.
Apply the latest techniques in parallelization and partitioning to automate kernel generation and exploit highly optimized execution paths.
Rapidly prototype and data-drive exploration of new runtime ideas.
Benchmark and analyze the outputs produced by our optimizing compiler on target hardware.
Build tools to collect and analyze performance bottlenecks.
Work closely with our product team to understand the evolving needs of ML engineers and drive improvements in runtime architecture.
Requirements Essential Skills and Experience
BS degree in Computer Science, Computer Engineering, or equivalent practical experience.
4+ years of experience working with compilers or runtime systems.
Deep understanding of asynchronous and concurrent programming.
4+ years of experience with C/C++ (C++14 or newer).
Understanding of hardware architecture: vector vs. scalar registers and instructions, memory hierarchies.
Knowledge of operating system kernel development or hypervisor development.
Preferred Skills and Experience
Master's or PhD in Computer Science, Computer Engineering, or equivalent.
Experience developing or maintaining GPU compute libraries such as CUDA or ROCm.
Experience with GPU programming and optimization.
Background in high-performance computing (HPC).
Knowledge of deep learning frameworks such as PyTorch, JAX, or Triton.
Experience programming large compute clusters.
Why Join Lemurian Labs
Build the runtime that makes next-generation AI infrastructure actually go fast.
Work across the full stack - from hardware intrinsics to compiler output to distributed execution.
Join a team that approaches infrastructure as a canvas, not a constraint.
Competitive compensation including equity, medical/dental/vision, retirement savings, and wellness benefits.
Compensation depends on experience and geographic location and will be narrowed during the interview process. Additional benefits include equity, company bonus opportunities, medical, dental, and vision coverage, a retirement savings plan, and supplemental wellness benefits.