On-Device AI Inference Engineer

Hark

$200K — $450K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4-8+ years in performance-critical software optimization, especially on accelerators like GPUs and DSPs.
  • Proficient in C/C++ with experience in SIMD, custom kernels, and memory optimization.
  • Deep understanding of transformer models, including attention mechanisms and KV-cache behavior.
  • Ability to optimize compute, memory, and power budgets in software development.
  • Proven track record of implementing optimized models on constrained hardware.

Responsibilities

  • Write and enhance low-level kernels for transformer workloads on specific silicon targets.
  • Manage model residency, scheduling, and memory sharing for efficiency across concurrent tasks.
  • Profile real hardware to identify and address performance bottlenecks.
  • Reduce model precision for compliance with size, latency, and power budgets.
  • Ensure smooth execution of transformer workloads on new hardware accelerators.
  • Communicate deployment constraints back to model teams to inform architectural choices.

Benefits

  • Flexible work hours to support work-life balance.
  • Collaborative, small team environment promoting hands-on contributions.
  • Opportunity to work closely with hardware teams on cutting-edge technology.
  • Exposure to a variety of hardware accelerators in AI applications.
  • Engagement with the latest in embedded AI silicon technologies.
Full Job Description
About the Role

You'll make Hark's models run fast on the hardware we ship. That means writing the kernels, building the runtime paths, and profiling transformer workloads on DSPs, NPUs, and other constrained targets until they hit the latency and power budgets our devices are built around. This is hands-on systems work close to the metal, on a small team where the code you write is what users feel as response time.

Responsibilities
  • Write and optimize the low-level kernels and runtime paths that transformer workloads execute through on target silicon.
  • Decide how multiple models share limited memory and power i.e. residency, scheduling, and swap behavior across concurrent workloads.
  • Profile models on real hardware, find the bottlenecks, and close the gap between theoretical and delivered performance.
  • Take models from full precision to INT8/INT4 and get them running within per-product size, latency, and power budgets.
  • Get transformer workloads executing efficiently on new accelerators as they come online, working alongside the hardware team.
  • Feed real deployment constraints back to the model teams so architecture decisions account for what the hardware can actually do.

Requirements
  • 4-8+ years writing performance-critical software, with hands-on optimization on GPUs, NPUs, DSPs, or similar accelerators.
  • Strong C/C++ and comfort with SIMD, custom kernels, memory layout, and the profiling tools that go with them.
  • You understand attention, KV-cache behavior, and where transformer inference actually spends its time and memory bandwidth.
  • You reason in compute, memory, and power budgets, and you've optimized against them rather than around them.
  • You've had a model you optimized run in a product on constrained hardware.

Bonus Qualifications
  • Experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Familiarity with ONNX Runtime, TVM, MLIR, TensorRT, or similar inference and compiler toolchains.
  • Background in speech or audio inference, where latency is perceptible to the user.

Compensation

The US base salary range for this full-time position is between $200,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Similar Jobs

More Jobs at Hark

More Consumer Technology Jobs

Find similar On-Device AI Inference Engineer jobs: