Member of Technical Staff - Inference

RadixArk

$130K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in systems engineering or ML infrastructure
  • Expertise in large-scale inference systems for LLMs or generative models
  • Deep understanding of GPU architecture and its performance characteristics
  • Experience optimizing latency and throughput in production systems
  • Strong knowledge of distributed systems and networking basics
  • Proficient in C++, Rust, Go, or Python for production environments
  • Experience profiling and optimizing compute-intensive workloads

Responsibilities

  • Design and build large-scale inference systems for frontier AI models
  • Optimize latency, throughput, and GPU utilization in production inference
  • Develop and enhance model serving architectures and runtimes
  • Work on batching, scheduling, and memory management strategies
  • Collaborate across teams for performance optimization
  • Debug performance bottlenecks throughout the system stack
  • Drive reliability and scalability of inference infrastructure
  • Build tooling for observability and performance analysis
  • Contribute to long-term inference architecture and strategy

Benefits

  • Flexible work arrangements
  • Comprehensive benefits package
  • Equity options in the company
  • Opportunity to work with cutting-edge AI technologies
  • Collaboration with leading AI labs and cloud providers
Full Job Description
About the Role

RadixArk is seeking a Member of Technical Staff - Inference to push the limits of large-scale AI inference.

You will work on the core systems that serve frontier models at scale, optimizing performance, latency, throughput, and cost across thousands of GPUs. This role sits at the intersection of systems engineering, ML infrastructure, and performance optimization.

Your work will directly shape how state-of-the-art models are deployed and experienced by users worldwide.

This is a deeply technical, high-impact role for engineers who enjoy working close to the hardware-software boundary and solving performance-critical problems at scale.
Requirements
  • 5+ years of experience in systems engineering, ML infrastructure, or performance-critical backend systems
  • Strong expertise in large-scale inference systems for LLMs or generative models
  • Deep understanding of GPU architecture and performance characteristics
  • Experience optimizing latency- and throughput-critical production systems
  • Strong knowledge of distributed systems and networking fundamentals
  • Proficiency in Python, Rust, C++, or Go for production systems
  • Experience profiling and optimizing compute-intensive workloads
  • Strong debugging skills across system layers (model, runtime, kernel, network)
Strong Plus
  • Experience with LLM serving stacks (SGLang, vLLM, TensorRT-LLM, etc.)
  • Open-source contributions in ML or systems infrastructure
  • Familiarity with CUDA, Triton, or custom kernel optimization
  • Experience with batching, KV-cache management, and scheduling strategies
  • Experience running inference at scale (1000+ GPUs)
  • Background in HPC or high-performance systems
Responsibilities
  • Design and build large-scale inference systems for frontier AI models
  • Optimize latency, throughput, and GPU utilization in production inference
  • Develop and improve model serving architectures and runtimes
  • Work on batching, scheduling, and memory management strategies
  • Collaborate with kernel, compiler, and systems teams on performance optimization
  • Debug performance bottlenecks across the stack
  • Drive reliability and scalability of inference infrastructure
  • Build tooling for observability, profiling, and performance analysis
  • Contribute to long-term inference architecture and strategy
Compensation

We offer competitive compensation with meaningful equity, comprehensive benefits, and flexible work arrangements. Compensation depends on location, experience, and level.

Similar Jobs

More Jobs at RadixArk

  • Member of Technical Staff - Inference
    $130K — $180K *
    Palo Alto, CA 94303 (Santa Clara County)
    Information Technology
    In-Person
  • Product Manager
    $120K — $150K *
    Palo Alto, CA 94303 (Santa Clara County)
    Information Technology
    In-Person
  • Backend/Platform Engineer
    $120K — $160K *
    Palo Alto, CA 94303 (Santa Clara County)
    Information Technology
    In-Person
  • Designer
    $80K — $120K *
    Palo Alto, CA 94303 (Santa Clara County)
    Consumer Technology
    In-Person
  • Product Marketing Manager
    $120K — $150K *
    Palo Alto, CA 94303 (Santa Clara County)
    Enterprise Technology
    In-Person

More Information Technology Jobs

Find similar Member of Technical Staff - Inference jobs: