Working onsite at our Santa Clara, CA, headquarters 3 days per week hybrid. Will consider remote in the United States. Will consider remote in the United States.
The role: Principal Software Engineer, Performance Analysis and Modelingd-Matrix is looking for a computer engineer to help analyze and model performance across the hardware/software boundary of our AI inference accelerators. This role focuses on emerging hardware technologies (DIMC, D2D, 3D-DRAM) and emerging workloads (generative inference, multi-modal LLMs) - building the analytical models and simulation tools that let the architecture team project performance on current and future d-Matrix silicon. You'll work closely with Hardware Design, Compiler, Inference Server, and Kernels teams, translating workload analysis into concrete modeling inputs and surfacing HW/SW improvement opportunities. This is a hands-on technical contributor role within the broader architecture organization.
What You Will Do- Analyze emerging ML workloads, multi-modal LLMs, CoT reasoning models, and video/audio generation to identify performance-relevant properties.
- Build and maintain analytical performance models that project behavior on current and future d-Matrix hardware generations.
- Develop and extend architecture simulators to support performance analysis of proposed HW/SW features.
- Partner with hardware design, compiler, inference server, kernel, and product teams to validate modeling assumptions and surface downstream implications.
- Track relevant ML architecture and algorithms research and incorporate findings into modeling work.
- Propose targeted HW/SW feature improvements based on modeling results and workload analysis.
- Document modeling methodology and findings for reuse across the architecture team.
What You Will Bring- BSEE with 6+ years of industry experience, or MSEE with 4+ years of industry experience.
- Working knowledge of computer architecture, HW/SW co-design, performance modeling, and ML fundamentals (particularly DNNs).
- Programming fluency in C/C++ or Python.
- Experience building or working with analytical performance models or architecture simulators.
- Self-motivated and collaborative, comfortable working across hardware and software teams.
Preferred Qualifications- Experience optimizing AI/ML workloads on accelerator technologies
- Research and investigation in AI/ML architecture/microarchitecture