Technical Lead, On-Device AI Inference

Hark

$300K — $500K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8-12+ years in high-performance computing with a focus on GPUs, NPUs, or specialized accelerators
  • Deep understanding of attention mechanisms, quantization, and memory bandwidth
  • Experience designing or optimizing inference engines and writing custom kernels
  • Proven track record leading teams on performance-critical software
  • Experience bringing models from research to real-world deployment on constrained hardware

Responsibilities

  • Evaluate and recommend GPUs, NPUs, DSPs, and other accelerators for deployment
  • Collaborate with ML teams to shape deployment architectures
  • Develop low-level execution layers and runtime systems for transformer workloads
  • Engage with silicon vendors to enable efficient transformer execution
  • Lead a team of engineers to set standards for the inference stack

Benefits

  • Comprehensive benefits package including health, dental, and vision insurance
  • 401(k) retirement plan with company match
  • Flexible work arrangements including remote options
  • Generous paid time off and parental leave
  • Professional development and continuing education support
Full Job Description
About the Role

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

Responsibilities
  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment, and own the recommendation hardware decisions are made against.
  • Work with the foundation model and audio ML teams to shape architectures that meet deployment constraints before training locks them in.
  • Build the low-level execution layer, custom kernels, runtime systems, and compiler paths that transformer workloads run through on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up new accelerators and get efficient transformer execution on them early.
  • Hire and lead a team of engineers on performance-critical software, and set the technical bar for the inference stack.

Requirements
  • 8-12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
  • Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits.
  • You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters.
  • Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one.
  • You've taken a model from a research checkpoint to running on constrained hardware in a product people use.

Bonus Qualifications
  • Hands-on experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Experience with speech, audio, or streaming multimodal inference where latency is perceptible to the user.
  • Contributions to open-source inference or compiler toolchains (TensorRT, ONNX Runtime, TVM, MLIR, and similar).

Compensation

The US base salary range for this full-time position is between $300,000 - $500,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Similar Jobs

More Jobs at Hark

More Enterprise Technology Jobs

Find similar Technical Lead, On-Device AI Inference jobs: