Staff AI Scheduling & Orchestration Engineer

Bitdeer Technologies Group

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's in Computer Science, Electrical Engineering, or related field
  • 6+ years in distributed systems engineering with Kubernetes expertise
  • Hands-on experience with AI workload execution patterns
  • Proven ability to scale scheduling stacks in HPC or cloud environments
  • Strong knowledge of GPU architectures and associated scheduling challenges
  • Familiarity with infrastructure automation tools like Terraform
  • Excellent technical communication and leadership skills

Responsibilities

  • Design and implement advanced batch scheduling architectures for AI workloads
  • Develop cluster-wide admission control and sophisticated job queueing mechanisms
  • Leverage Kubernetes Dynamic Resource Allocation for complex requests
  • Architect topology-aware pod placement strategies for low-latency communication
  • Implement automated GPU sharing technologies and multi-tenancy policies
  • Collaborate with hardware and storage teams for tight integration
  • Drive reliability and scalability of the scheduling stack

Benefits

  • Opportunities for mentorship and professional growth
  • Exposure to cutting-edge technology in AI and cloud environments
  • Collaboration with cross-functional teams in a high-velocity context
  • Opportunity to influence architectural decisions and solutions
  • Work on large-scale distributed training and inference systems
Full Job Description
Position Overview

We are seeking a Staff AI Scheduling & Orchestration Engineer to lead the workload placement logic that defines our AI-native NeoCloud platform. Standard Kubernetes scheduling is insufficient for the demands of large-scale AI; you will be responsible for eliminating "GPU stranding" and maximizing utilization across our expensive compute fleets. This role is pivotal in building a high-performance scheduling fabric that understands the physical realities of our hardware-from NVLink-connected GPU topologies to InfiniBand interconnects. You will work at the intersection of distributed systems and AI, driving the architectural decisions that enable our platform to handle massive-scale distributed training and inference jobs with industry-leading efficiency.

Key Responsibilities
  • Design and implement advanced batch scheduling architectures using frameworks like Volcano or YuniKorn to support multi-node gang scheduling.
  • Develop and manage cluster-wide admission control and sophisticated job queueing mechanisms utilizing Kueue to manage high-volume AI workload traffic.
  • Leverage Kubernetes Dynamic Resource Allocation (DRA) and custom scheduler plugins to manage complex accelerator requests natively.
  • Architect topology-aware pod placement strategies that optimize for low-latency communication via NVLink and InfiniBand fabrics.
  • Implement automated GPU sharing technologies (e.g., MIG, time-slicing) and multi-tenancy isolation policies to maximize cluster-wide utilization.
  • Collaborate with the GPU Systems and Storage teams to ensure the scheduling layer is tightly integrated with bare-metal hardware and storage I/O patterns.
  • Drive the reliability and scalability of the scheduling stack, resolving resource contention and deadlock scenarios in large-scale HPC environments.
  • Mentor junior engineers and conduct design reviews to maintain architectural excellence in our orchestration layer.

Qualifications
  • Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.
  • 6+ years of distributed systems engineering, with deep, hands-on expertise in Kubernetes scheduling frameworks and orchestrators.
  • Extensive experience with AI workload execution patterns and distributed training frameworks (e.g., PyTorch Distributed, Ray, MPI).
  • Proven track record of operating, debugging, and scaling scheduling stacks in high-performance computing (HPC) or large-scale production cloud environments.
  • Strong knowledge of GPU hardware architectures and the specific scheduling challenges related to distributed AI training and inference.
  • Experience with infrastructure automation and infrastructure-as-code (e.g., Terraform, Go-based Operators).
  • Excellent technical communication and leadership skills; ability to influence cross-functional teams and align architectural goals.
  • Ability to work in a high-velocity engineering environment and translate complex, ambiguous requirements into concrete, scalable engineering solutions.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar Staff AI Scheduling & Orchestration Engineer jobs: