Member of Technical Staff, Inference

Reactor

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in a relevant technical field or equivalent experience
  • Foundational skills in systems programming with experience in bottleneck resolution
  • Expertise in ML frameworks such as PyTorch and TensorRT
  • Proficient in model compilation and quantization techniques
  • Familiarity with GPU hardware, specifically NVIDIA
  • Strong grasp of transformer architectures and ML optimization methods

Responsibilities

  • Drive improvements in model performance specifically for diffusion models
  • Design and implement an in-house, high-performance inference runtime
  • Optimize performance using torch.compile and custom CUDA kernels
  • Enhance model efficiency through quantization and architectural changes
  • Profile and benchmark models to identify bottleneck issues
  • Collaborate with partner teams to integrate models into the platform

Benefits

  • Competitive salary and early equity opportunities
  • Visa sponsorship and relocation assistance
  • Comprehensive health, dental, and vision plans
Full Job Description
Member of Technical Staff, Inference

Department: Engineering

Employment Type: Full Time

Location: San Francisco

Description

We're looking for an ML Inference Engineer with deep expertise in high-performance ML engineering. This is a highly technical, high-impact role focused on squeezing every drop of performance from generative media models.

You'll work across the inference stack, designing novel frameworks, optimizing inference performance, and shaping Reactor's competitive edge in ultra-low-latency, high-throughput environments.

What You'll Do
  • Drive our frontier position on model performance for diffusion models
  • Design and implement a high-performance in-house inference runtime
  • Implement optimizations using torch.compile, custom CUDA kernels, and specialized inference frameworks
  • Optimize neural network models through quantization, pruning, and architectural modifications
  • Profile and benchmark model performance to identify computational bottlenecks
  • Collaborate directly with model partner teams to integrate their models into our platform


Required Skills
  • Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field (or equivalent practical experience)
  • Strong foundation in systems programming, with a track record of identifying and resolving bottlenecks
  • Deep expertise in PyTorch, TensorRT, TransformerEngine, Nsight, ONNX Runtime
  • Model compilation, quantization (INT8/FP16), and advanced serving architectures
  • Working knowledge of GPU hardware (NVIDIA)
  • Strong understanding of transformer architectures and modern ML optimization techniques


Benefits
  • Competitive SF salary and meaningful early equity
  • Visa sponsorship and relocation support
  • Generous health, dental, and vision coverage

Similar Jobs

More Jobs at Reactor

More Information Technology Jobs

Find similar Member of Technical Staff, Inference jobs: