Senior Researcher - Edge AI Optimization/Hardware-Aware ML

Huawei Technologies Canada Co., Ltd.

$110K — $130K *
Consumer Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • PhD or equivalent in Machine Learning, Computer Science, Electrical/Computer Engineering, or related field
  • Strong programming skills in Python and C/C++
  • Experience with AI toolchains and deep learning frameworks like PyTorch or TensorFlow
  • Proven publication record at top venues or patents in ML efficiency
  • 2+ years of relevant research or industry experience
  • Familiarity with optimization for edge hardware and understanding of power/thermal constraints
  • Technical skills in quantization, sparsity, compiler techniques, and memory management

Responsibilities

  • Conduct research in hardware-aware neural network optimization including techniques like quantization and pruning
  • Prototype and evaluate techniques for efficient inference under device constraints
  • Publish findings in academic and industry forums such as journals and workshops
  • Collaborate on compiler and runtime improvements for ML frameworks
  • Profile and optimize models using real device performance data
  • Build benchmarking methodologies for edge device targets
  • Drive optimization for on-device machine learning updates and personalization

Benefits

  • Opportunity to work on cutting-edge research in hardware-aware optimization
  • Collaboration with cross-functional teams including product engineering and hardware
  • Mentorship opportunities for junior researchers and engineers
  • Expose work to both academic and industry audiences through publications and presentations
  • Participation in defining technical roadmap for future innovations
Full Job Description
Huawei Canada has an immediate permanent opening for a Researcher.

About the job:
  • Conduct research in hardware-aware neural network optimization (e.g., quantization-aware training, mixed precision, pruning, distillation, neural architecture search).
  • Develop novel approaches for latency/energy-aware training objectives and Pareto optimization (accuracy vs. compute vs. memory).
  • Prototype and evaluate techniques for efficient inference under device constraints (thermal limits, memory bandwidth, intermittent connectivity).
  • Publish and present findings internally and externally (papers, workshops, patents, technical blogs).
  • Optimize inference pipelines across pre/post-processing, scheduling, operator fusion, memory planning, and runtime execution.
  • Collaborate on or contribute to compilers / runtimes (e.g., TVM, MLIR, XLA, TensorRT, ONNX Runtime, TFLite, ExecuTorch) to improve operator coverage and performance.
  • Profile and optimize models with real device traces, addressing bottlenecks such as cache misses, memory bandwidth, kernel launch overhead, and CPU-NPU handoff.
  • Build and maintain hardware-aware benchmarking methodology and regression suites for edge targets (ARM CPU, mobile GPU, DSP, NPU).
  • Create deployment recipes for heterogeneous compute (CPU+GPU+NPU) including partitioning strategies and fallback paths.
  • Drive optimization for on-device personalization and incremental updates when needed (e.g., small adapters, efficient fine-tuning).
  • Partner with product engineering, platform teams, and hardware teams to translate device constraints into research targets and to transition research prototypes into production.
  • Mentor junior researchers/engineers, review experimental designs, and raise the quality bar for measurement rigor and reproducibility.
  • Define technical roadmap areas (e.g., next-gen quantization, kernel optimization, model families for edge, compiler improvements).


About the ideal candidate:
  • PhD (or equivalent research experience) in Machine Learning, Computer Science, Electrical/Computer Engineering, or related field.
  • Strong programming skills in Python and C/C++ (or equivalent systems language).Experience building AI agent / harness / skill toolchains, including model evaluation, orchestration, and LLM-powered tooling. Hands-on experience with deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and deployment toolchains (e.g., ONNX, TFLite, TensorRT, TVM, MLIR-based stacks). Solid knowledge of performance profiling: latency measurement, memory profiling, kernel-level bottleneck analysis, and experimental rigor.
  • Proven publication record at top venues (e.g., NeurIPS/ICML/ICLR, MLSys, ASPLOS, ISCA, MICRO) and/or patents in ML efficiency.
  • 2+ years of relevant experience (research lab or industry) with demonstrated impact in at least one of:
    • Model compression (quantization/pruning/distillation)
    • Efficient architectures (MobileNet-like, MoE, efficient transformers, etc.)
    • ML systems/compilers/runtime optimization
    • hardware-aware optimization for edge deployment
  • Experience optimizing for specific edge hardware:
    • ARM NEON, mobile GPUs, DSPs, NPUs, microcontrollers
    • Experience with distributed benchmarking, CI for performance regression, and reproducible experiment pipelines.
    • Understanding of power/thermal constraints and methodologies for measuring energy on device.
    • Experience with efficient LLM/VLM inference on edge (KV-cache optimization, quantized attention, speculative decoding, etc.).
  • Technical Skills:
    • Quantization: PTQ/QAT, per-channel/per-tensor, calibration, smooth quant, GPTQ-like methods, mixed precision
    • Sparsity: structured pruning, N: M sparsity, hardware-friendly sparsity
    • Compiler techniques: graph rewriting, operator lowering, scheduling, kernel autotuning
    • Runtime techniques: memory arenas, tensor lifetime analysis, static vs dynamic shapes, batching strategies
    • Hardware fundamentals: cache hierarchy, SIMD, memory bandwidth, accelerator programming models

Additional Information:

Similar Jobs

More Jobs at Huawei Technologies Canada Co., Ltd.

More Consumer Technology Jobs

Find similar Senior Researcher - Edge AI Optimization/Hardware-Aware ML jobs: