BSEE with 10+ years of industry experience or MSEE with 8+ years.
Expertise in computer architecture, hardware/software co-design, and performance modeling.
Strong fundamentals in machine learning, particularly deep neural networks (DNNs).
Proficiency in programming languages C/C++ and Python.
Experience in analytical performance models and architecture simulators.
Publication record in top-tier architecture or machine learning conferences is a significant advantage.
Responsibilities
Analyze the latest machine learning workloads, including multi-modal LLMs and reasoning models.
Develop hardware and software features for next-generation inference accelerators in datacenters.
Stay updated with the latest research in ML architecture and algorithms.
Collaborate with cross-functional teams including Product, Hardware design, and Compiler.
Create analytical models to project performance on current and future hardware.
Identify performance implications of emerging ML algorithms and workloads.
Propose new hardware/software features to improve algorithm performance.
Benefits
Flexible hybrid work model with onsite presence 3 days a week.
Inclusive team culture that values diverse perspectives.
Focus on personal and professional growth through challenges.
Direct communication and collaborative approach in the workplace.
Full Job Description
Working onsite at our Santa Clara, CA headquarters 3 days per week Hybrid.
The role: Principal Architect- Performance Analysis and Modeling
d-Matrix is seeking outstanding computer architects to help accelerate AI application performance at the intersection of both hardware and software, with particular focus on emerging hardware technologies (such as DIMC, D2D, 3D-DRAM etc.) and emerging workloads (such as generative inference etc.). Our acceleration philosophy cuts through the system ranging from efficient tensor cores, storage, and data movements along with co-design of dataflow, and collective communication techniques.
What you will do:
As a member of the architecture team, you will analyze the latest ML workloads (multi-modal LLMs, CoT reasoning models, video/audio-generation)
You will contribute Hardware and Software features that power the next generation of inference accelerators in datacenters.
This role requires to keep up the latest research in ML Architecture and Algorithms, and collaborate with different partner teams including Product, Hardware design, Compiler, Inference Server, Kernels.
Your day-to-day work will include (1) analyzing the properties of emerging machine learning algorithms and workloads and identifying functional, performance implications (2) Creating analytical models to project performance on current and future generations of d-matrix hardware (3) proposing new HW/SW features to enable or accelerate these algorithms
What you will bring:
Minimum:
BSEE with 10+ years of industry experience or MSEE preferred with 8+ years of industry experience.
Solid grasp through academic or industry experience in multiple of the relevant areas - computer architecture, hardware software codesign, performance modeling, ML fundamentals (particularly DNNs).
Programming fluency in C/C++ or Python.
Experience with developing analytical performance models, architecture simulators for performance analysis,
Research background with publication record in top-tier architecture, or machine learning venues is a huge plus (such as ISCA, MICRO, ASPLOS, HPCA, DAC, MLSys etc.).
Self-motivated team player with strong sense of collaboration and initiative.