Quadric.io

AI Performance Modeling Engineer

Quadric.io$150K — $200K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong Python skills for quantitative modeling and technical validation.
  • Solid understanding of computer architecture principles, including memory hierarchies and execution bottlenecks.
  • Proven technical writing ability for clear and evidence-based presentations of conclusions.
  • Deep knowledge of neural network inference operators and tensor shapes, or extensive quantitative modeling experience.
  • Degree in Computer Science, Electrical Engineering, or Computer Engineering, or equivalent practical experience.

Responsibilities

  • Build cycle-level performance models for AI inference workloads on GPNPU architecture.
  • Derive hardware lane operations and software pipelining overlaps from first principles.
  • Model tensor placement and data movement across various memory tiers.
  • Adapt workloads for vision networks and LLMs, incorporating architectural details.
  • Calibrate models against simulators and profiling data to achieve accuracy.
  • Draft technical studies that influence architecture and product direction.
  • Balance latency and throughput performance in model evaluations.

Benefits

  • Medical, dental, and vision insurance with 99% premium coverage for employees.
  • Company-paid life insurance and optional supplemental coverage.
  • Short-term and long-term disability insurance.
  • Commuter support, including parking or Caltrain reimbursements.
  • Flexible Spending Account (FSA) and Health Savings Account (HSA) options.
  • Equity ownership in the company and eligibility for performance bonuses.
  • Paid parental leave and a 401(k) retirement plan.
  • Flexible Paid Time Off (PTO) and winter holiday shutdown.
  • Daily catered lunches and a collaborative office environment in a prime location.
Full Job Description
The Opportunity

Quadric has created an innovative General-Purpose Neural Processing Unit (GPNPU) architecture. Unlike standard accelerators, the Quadric GPNPU executes both neural network graph code and conventional C++ DSP/control code across edge and endpoint devices.

As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels-from early-career engineers to seasoned experts-with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware.
What Youll Do
Performance Modeling & Architectural Analysis
  • Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware.
  • Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps.
  • Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers.
  • Model sharding and collective boundary communication across multi-die systems.
Workload Adaptation & Technical Writing
  • Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats.
  • Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets.
  • Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions.
  • Balance single-stream latency against scaled throughput performance.
What Success Looks Like

Within your first 6-12 months, youll:
  • Own a full workloads model end to end, calibrated against simulation and trusted by the engineering team.
  • Build performance models that consistently predict workload behavior within 10-15% of actual measurements.
  • Publish a written study whose defended conclusions directly shape an architecture or product decision.
  • Review and extend performance models beyond your initial starting domain.
What Were Looking For
Required
  • Python & Quantitative Modeling: Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code.
  • Computer Architecture Fundamentals: Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks (via industry experience, coursework, or research).
  • Technical Writing: Comfort writing clear technical studies that state and defend evidence-based conclusions.
  • Core Technical Depth (One of the following):
    • Option A: Deep understanding of NN inference operators and tensor shapes (e.g., Transformers, attention mechanisms, MoE, prefill/decode split).
    • Option B: Proven performance modeling experience in another quantitative/technical domain.
  • Education: BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
Preferred
  • Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels.
  • Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim).
  • Background in compiler internals (cost models, autotuners) or proficiency in C++.
  • Published performance studies or technical write-ups.
What We Offer

The base salary range for this position is $150,000 to $200,000. This range reflects the full span of levels and geographies at which Quadric hires for this role. The actual base salary offered will depend on a number of factors, including the specific level of the role, years and depth of relevant experience, technical skills and competencies, the criticality of the role to the business, internal equity, and work location. In addition to base salary, this role is eligible for equity and a discretionary annual performance bonus as applicable to the role and level.

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance from day one - Premiums covered at 99% for Employees
  • Company-paid life Insurance
  • Voluntary supplemental life insurance
  • STD + LTD insurance
  • Commuter support including parking or Caltrain reimbursement. Our office is conveniently located within walking distance of the Caltrain station
  • FSA + HSA
  • Equity with the business
  • Paid Parental Leave
  • 401(k) Retirement Plan
  • Flexible PTO
  • Winter holiday shutdown
  • Catered lunch each day in our office
  • Downtown Burlingame office location, close to shops, cafes, and local amenities
  • Collaborative, low-ego culture with significant ownership and impact
  • A work culture focused on innovative disruption

About Quadric.io

Quadric.io is a computer hardware company that specializes in developing high-performance processors for artificial intelligence applications. The company was founded in 2016 and is based in Palo Alto, California. Quadric.io's processors are designed to be highly efficient and scalable, making them ideal for use in data centers and other large-scale computing environments. The company has received significant investment from venture capital firms and has attracted attention for its innovative approach to processor design.
Learn more about Quadric.io
Size
50 employees
Industry

Similar Jobs

More Jobs at Quadric.io

More Information Technology Jobs

Find similar AI Performance Modeling Engineer jobs: