Position OverviewWe are seeking a Staff AI Scheduling & Orchestration Engineer to lead the workload placement logic that defines our AI-native NeoCloud platform. Standard Kubernetes scheduling is insufficient for the demands of large-scale AI; you will be responsible for eliminating "GPU stranding" and maximizing utilization across our expensive compute fleets. This role is pivotal in building a high-performance scheduling fabric that understands the physical realities of our hardware-from NVLink-connected GPU topologies to InfiniBand interconnects. You will work at the intersection of distributed systems and AI, driving the architectural decisions that enable our platform to handle massive-scale distributed training and inference jobs with industry-leading efficiency.
Key Responsibilities- Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
- Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
- Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
- Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
- Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
- Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
- Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
- Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.
Qualifications- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
- 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
- Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
- Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
- Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
- Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
- Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
- Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.