About the Role You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.
Responsibilities- Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment, and own the recommendation hardware decisions are made against.
- Work with the foundation model and audio ML teams to shape architectures that meet deployment constraints before training locks them in.
- Build the low-level execution layer, custom kernels, runtime systems, and compiler paths that transformer workloads run through on target hardware.
- Partner with silicon vendors and internal hardware teams to bring up new accelerators and get efficient transformer execution on them early.
- Hire and lead a team of engineers on performance-critical software, and set the technical bar for the inference stack.
Requirements- 8-12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
- Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits.
- You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters.
- Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one.
- You've taken a model from a research checkpoint to running on constrained hardware in a product people use.
Bonus Qualifications- Hands-on experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
- Experience with speech, audio, or streaming multimodal inference where latency is perceptible to the user.
- Contributions to open-source inference or compiler toolchains (TensorRT, ONNX Runtime, TVM, MLIR, and similar).
CompensationThe US base salary range for this full-time position is between $300,000 - $500,000 annually.
The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.