About the RoleAs a Staff GPU Inference SDET, you will be the founding quality, reliability, and validation lead for a new GPU Inference Development team. Working closely with engineering leads and cross-functional systems infrastructure teams, you will design, build, and scale the end-to-end release qualification and automated test ecosystem for our GPU inference stack and rack-scale accelerated compute fleets. In this high-impact role, you will be responsible for building automated test suites to validate multi-node GPU cluster bring-up, verifying prefill worker optimizations, testing open-source and custom serving engines, and ensuring numerical correctness and performance stability under real-world streaming workloads. You will be the primary technical anchor ensuring production-grade reliability, fault isolation, and peak inference performance across accelerated GPU infrastructure.
WHAT YOU'LL DOBuild GPU Release Qualification Systems:
Design and implement automated test automation frameworks, regression gates, and release qualification pipelines for the complete GPU inference stack-spanning custom API services, model-serving workers, container runtimes, serving engines, driver stacks, and firmware.
Inference Serving & Workload Validation:
Benchmark and stress-test distributed LLM serving frameworks, focusing on prefill vs. decode worker performance, continuous batching, prefix caching, KV-cache efficiency, and tensor/expert parallelism.
Performance & Performance Modeling Verification:
Build automated workload replay and benchmarking tools to validate GPU performance models. Track critical serving metrics including Time-to-First-Token (TTFT), Inter-Token Latency (ITL), request throughput, tail latency (P99), and capacity efficiency.
Numerical Correctness & Quality Gates:
Build validation infrastructure to ensure model accuracy, precision stability (FP16/FP8/quantization), determinism, and output correctness across software updates, kernel fusions, and hardware revisions.
Fault Injection & Fleet Resilience:
Engineer chaos engineering and fault-injection suites to simulate node failures, inter-node network degradation, GPU memory leaks, driver/firmware mismatches, and automated recovery paths for multi-node GPU clusters.
Observability & CI/CD Integration:
Integrate automated test pipelines with telemetry tools (e.g., Prometheus, Grafana) to turn one-off investigations into repeatable engineering gates and continuous performance monitoring.
REQUIREMENTS:8+ years of software engineering experience as an SDET, Infrastructure Quality Lead, or Systems Test Engineer.
GPU & Cluster Infrastructure Expertise:
Hands-on experience bringing up, provisioning, and validating multi-node GPU clusters (NVIDIA or AMD ecosystem) across public cloud infrastructure or enterprise data center environments.
Inference Stack Knowledge:
Deep understanding of LLM serving engines and distributed runtimes, including prefill vs. decode disaggregation, KV-cache management, and dynamic batching.
Automation & Scripting:
Expert-level Python programming skills with extensive experience designing custom test automation frameworks, diagnostic tooling, and CI/CD integration.
Orchestration & Networking:
Strong proficiency with container orchestration tools (e.g., Kubernetes, Slurm, Ray) and high-performance cluster interconnects (e.g., InfiniBand, RoCE, NCCL).
Failure Analysis & Debugging:
Proven background in root-cause analysis across software/hardware boundaries, stress testing, and node failure simulation in distributed systems.
NICE TO HAVES: - Direct experience with either AMD (ROCm / HIP) or NVIDIA software stacks.
- Experience building workload replay tools, ML evaluation pipelines, or MLPerf Inference benchmark suites.
- Familiarity with low-level kernel profiling tools (PyTorch Profiler, NVTX, ROCm profilers) or C++
Apply today and become part of the forefront of groundbreaking advancements in AI!