Staff Modeling Architect

Neurophos Inc

$160K — $190K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent experience.
  • 8+ years in hardware modeling or performance analysis for architecture, RTL, compiler, or silicon teams.
  • Proven record of delivering a model or study relied upon by other teams.
  • Ability to select appropriate analytical methods for performance modeling.
  • Strong foundation in computer architecture and AI accelerators.
  • Proficiency in modern C++ (C++17 or later) for various models and simulations.
  • Experience with Python libraries for modeling, including NumPy and Matplotlib.

Responsibilities

  • Bring up and analyze inference workloads including transformers and quantization models.
  • Bind workloads from Hugging Face and PyTorch to the programming model for functional execution.
  • Co-design architectural aspects like tiling, scheduling, and memory hierarchy.
  • Conduct performance analysis to resolve both hardware and compiler bottlenecks.
  • Develop energy and latency models using Python for various traffic and computation scenarios.
  • Implement bit-accurate functional models for key hardware components.
  • Contribute to development of a cycle-accurate simulation kernel and performance models.

Benefits

  • Comprehensive health insurance plans.
  • Flexible work schedules with the possibility of remote options.
  • Professional development opportunities and mentoring programs.
  • Wellness programs and activities to support work-life balance.
  • Generous time off policy including paid holidays.
Full Job Description
Location: Austin, TX or Sunnyvale, CA. Full-time onsite position.

Reports To: Sr. Director of Modeling

FLSA Status: Exempt

Position Overview

We are seeking a staff-level modeling architect to build the path from a production model or application to two things: a performance and energy number Neurophos will stand behind, and a functional model that software can boot against before tape-out.

The T100 architecture is still moving, and the workloads are the models the industry is publishing now, so this is hardware/software co-design in practice. You will bind a workload to the programming model and runtime, run it on the model stack, and feed the result back into decisions on tiling, instruction set architecture (ISA), memory hierarchy, and multi-chip mapping. The team works between the principal architects, the RTL and physical design groups, and the compiler and runtime teams. At this level, you own a workload or block area along with the methodology behind it, and you mentor the engineers building models in that area.

Key Responsibilities
  • Bring up inference workloads as they ship, including dense and Mixture of Experts (MoE) transformers, attention and KV cache, expert routing, quantization, and hybrid/SSM models, plus retrieval, speech, vision, and recommendation workloads where they map onto the accelerator.
  • Bind Hugging Face and PyTorch workloads to the programming model and runtime, then run them on the functional model so that software and architecture are looking at the same behavior.
  • Co-design tiling, scheduling, the instruction set architecture (ISA), the SRAM and High Bandwidth Memory (HBM) hierarchy, network-on-chip (NoC) traffic, and multi-chip mapping across pipeline, tensor, and sequence parallelism, including collectives.
  • Run roofline and limiter analysis and design space exploration across microarchitecture options, resolving bottlenecks between the compiler view and the hardware.
  • Develop Python energy and latency models in NumPy, Pandas, and Matplotlib that cover operators, tiling, SRAM and HBM traffic, and optical GEMM and vector-unit time.
  • Implement bit-accurate C++ functional models of optical GEMM, SRAM vector processors, dataflow engines, and HBM, including narrow arithmetic, so software can begin bring-up before tape-out.
  • Contribute to the C++ event-driven simulation kernel itself, including coroutines, timed components, and traces, rather than only calling into it.
  • Implement cycle-approximate and cycle-accurate performance, power, and area (PPA) models, and align them with RTL through Verilator, SystemVerilog, and co-simulation.
  • Keep numbers consistent across roofline, limiter, performance model, and RTL simulation of the same workload, and document where they disagree.
  • Set the modeling methodology for a workload area, deciding what gets modeled at which fidelity and how to arbitrate when models disagree.
  • Maintain the interface and register specs as the source of truth for generating the C++ and SystemVerilog views, and mentor the engineers building models in your area.


Qualifications
  • BS, MS, or PhD in Computer Engineering, Electrical Engineering, Computer Science, or equivalent practical experience.
  • 8+ years of experience in hardware modeling, functional modeling, performance modeling, performance simulation, or accelerator performance analysis used by architects, RTL, compiler and runtime, or silicon teams. Graduate research may count toward this.
  • Track record of shipping a model or study that another team depended on, whether architecture, compiler, customer, or silicon.
  • Judgment to pick the right method for a given question among roofline, limiter analysis, analytical performance models, trace-driven simulation, transaction-level modeling (TLM), and RTL simulation.
  • Strong grounding in computer architecture, microarchitecture, memory systems, and AI accelerators, whether GPU, TPU, NPU, or custom SoC.
  • Modern C++ (C++17 or later) for functional models, performance models, and simulation infrastructure.
  • Python for models, analysis, and plots, including NumPy, Pandas, and Matplotlib.
  • Experience working inside a discrete-event, cycle-approximate, or cycle-accurate simulator such as SystemC, gem5, SST, or a custom kernel, rather than only driving one.
  • Ability to build an LLM or accelerator workload from a model card or paper, covering prefill and decode, MoE, GEMM tiling, and quantization.
Preferred Skills
  • PhD in Computer Engineering, Electrical Engineering, or Computer Science.
  • Hardware/software co-design alongside compiler, runtime, or ISA work, including MLIR, TVM, XLA, ONNX, operator fusion, or graph compilers.
  • Experience modifying or extending a simulation kernel, or correlating an analytical model against silicon, vendor datasheets, or measured datacenter GPUs and inference accelerators.
  • Familiarity with TLM 2.x, Verilator, SystemVerilog, DPI, or UVM.
  • Familiarity with HBM, DRAM controllers, cache, SRAM, network-on-chip (NoC), AXI, DMA, and scratchpad memory.
  • Power modeling with McPAT, CACTI, or a custom flow, plus FPGA prototyping or hardware emulation.

Similar Jobs

More Jobs at Neurophos Inc

More Enterprise Technology Jobs

Find similar Staff Modeling Architect jobs: