DRW

AI Inference Platform Engineer

DRW • $200K — $250K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience with large language models on NVIDIA GPUs
  • Deep understanding of modern inference runtimes like TensorRT-LLM or vLLM
  • Proficient in performance optimization techniques including speculative decoding and quantization
  • Strong knowledge of KV cache architecture and multi-tier caching
  • Experience with multi-tenant GPU scheduling and workload isolation
  • Hands-on experience with observability tools like Nsight and Prometheus
  • Solid experience in model serving infrastructure, including CI/CD practices

Responsibilities

  • Optimize LLM inference performance on NVIDIA GPU architectures
  • Build performance profiling and observability for multi-node inference systems
  • Design distributed inference architectures and optimize KV cache
  • Oversee model onboarding, including serving configuration analysis
  • Maintain performance profiles for hardware and model combinations
  • Monitor quality across different serving configurations
  • Manage the model lifecycle including versioning and rollback
  • Collaborate with SRE teams to automate deployment and improve operations
  • Optimize resource allocation and scheduling for GPU workloads

Benefits

  • Comprehensive medical, pharmacy, dental, and vision insurance
  • 401k plan with discretionary employer match
  • Short and long-term disability insurance
  • Life and AD&D insurance
  • Health savings and flexible spending accounts
  • Annual discretionary bonus eligibility
Full Job Description
About the Role

We're looking for an AI Inference Platform Engineer to build, operate, and optimize the systems that serve large language, vision, multimodal, and embedding models across DRW. This role provides DRW's firmwide interface to modern AI models, from early evaluation through reliable production use.

You'll work across inference runtimes, distributed systems, and production platform engineering, with deep GPU literacy. You'll own the serving platform end-to-end: onboarding newly released models, measuring quality and performance equivalence across serving configurations, scheduling workloads across tenants, and continuously improving latency, throughput, utilization, reliability, and cost across the inference fleet.

What You'll Do
  • Optimize LLM inference performance across modern NVIDIA GPU architectures and inference runtimes.
  • Build end-to-end performance profiling and observability to identify bottlenecks from individual GPU kernels through multi-node inference systems.
  • Design and optimize KV cache and distributed inference architectures, including caching, routing, memory tiering, and prefill/decode strategies.
  • Own day-0 model onboarding, determining the appropriate runtime, precision, sharding, memory, batching, cache policy, and serving configuration for new models.
  • Maintain validated performance profiles for important model and hardware combinations, including performance and quality regression testing.
  • Measure and monitor quality equivalence across serving configurations, including KV cache quantization, speculative decoding acceptance thresholds, precision choices, and model routing, so in-house serving can be trusted to match reference-model quality on production workloads.
  • Manage the production serving lifecycle of models, including versioning, compatibility, staging, canarying, promotion, rollback, and retirement.
  • Partner with SRE and platform teams to automate model deployment, distribution, production readiness, observability, and reliable operation across environments.
  • Optimize model placement, scaling, and resource allocation across the inference fleet to improve utilization and cost efficiency while meeting performance and reliability requirements.
  • Design and operate multi-tenant scheduling and isolation across shared GPU capacity, balancing latency SLOs, throughput, and priority across concurrent workloads.


What We're Looking For

The Tech
  • Hands-on experience serving LLMs on NVIDIA GPUs, with familiarity across current and emerging architectures (Hopper, Blackwell, and successors), HBM, Tensor Cores, NVLink/NVSwitch, and the compute and memory bottlenecks that shape serving decisions.
  • Deep expertise in at least one modern inference runtime such as TensorRT-LLM, vLLM, or SGLang.
  • Practical knowledge of inference optimization techniques including continuous batching, scheduling, chunked prefill, speculative decoding, quantization, CUDA Graphs, and paged attention.
  • Understanding of KV cache architecture, including prefix caching, block management, sizing, eviction, quantization, cache-aware routing, and multi-tier caching.
  • Experience measuring model quality equivalence across serving configurations, including evaluation harnesses, task-specific benchmarks, and regression detection for quantization, KV cache, and speculative decoding changes.
  • Experience designing and tuning distributed inference systems, including tensor parallelism, multi-node deployments, and disaggregated prefill and decode.
  • Experience with multi-tenant GPU scheduling, workload isolation, and QoS across concurrent inference workloads.
  • Proficiency with GPU performance and observability tooling such as Nsight, DCGM, OpenTelemetry, Prometheus, and Grafana.
  • Strong Linux and systems performance fundamentals, with the ability to diagnose bottlenecks across hardware, drivers, runtimes, networking, and application layers.
  • Production experience with model serving infrastructure, including CI/CD, automated testing, observability, and production readiness.

The Intangibles
  • You take a measurement-driven approach to performance optimization.
  • You take ownership of performance problems across hardware, runtime, model, and infrastructure boundaries.
  • You can move quickly and reprioritize as trading needs change, while maintaining a high bar for production systems.
  • You understand the importance of reliability, predictability, and performance when AI systems are integrated into trading workflows and decision-making processes.
  • You can evaluate unfamiliar models, runtimes, and hardware quickly and make sound engineering decisions with limited prior guidance.
  • You communicate clearly and can explain complex performance tradeoffs across engineering teams.


The annual base salary range for this position is $200,000 to $250,000 depending on the candidate's experience, qualifications, and relevant skill set. The position is also eligible for an annual discretionary bonus. In addition, DRW offers a comprehensive suite of employee benefits including group medical, pharmacy, dental and vision insurance, 401k (with discretionary employer match), short and long-term disability, life and AD&D insurance, health savings accounts, and flexible spending accounts.

[#LI-VD1]

About DRW

DRW is a financial trading firm that specializes in derivatives trading. The company was founded in 1992 by Don Wilson and has since grown to have offices in Chicago, London, Montreal, New York, and Singapore. DRW trades a variety of financial instruments, including futures, options, and cryptocurrencies. The company is known for its quantitative trading strategies and has developed a number of proprietary trading systems. DRW is also involved in venture capital and has invested in a number of technology startups.
Learn more about DRW
Size
1,000 employees
Industry
Founded
1992

Similar Jobs

More Jobs at DRW

More Information Technology Jobs

Find similar AI Inference Platform Engineer jobs: