Sr GPU Cloud K8S Expert (SRE SME)

Bitdeer Technologies Group

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Kubernetes operations, focusing on GPU workloads.
  • Deep expertise with Nvidia GPU operators in Kubernetes.
  • Hands-on experience in topology-aware scheduling.
  • Proficiency in Terraform, Helm, and GitOps workflows.
  • Strong background in site reliability engineering (SRE) principles.
  • Experience with monitoring tools like Prometheus and Grafana.
  • Programming skills in Go or Python for operator/CRD development.

Responsibilities

  • Design, deploy, and operate production Kubernetes clusters for GPU workloads.
  • Implement topology-aware scheduling for GPU and NVLink domains.
  • Manage custom resource definitions for GPU lifecycle.
  • Integrate AI frameworks such as Slurm and Kubeflow with Kubernetes.
  • Ensure multi-tenant isolation through namespaces and RBAC policies.
  • Automate bare-metal provisioning and lifecycle management for tenants.
  • Monitor cluster performance and automate incident management processes.

Benefits

  • Flexible working arrangements to promote work-life balance.
  • Opportunity to work with cutting-edge AI and GPU technologies.
  • Access to professional development resources and training.
  • Inclusive and collaborative team culture.
  • Participation in innovation-driven projects that impact real customers.
Full Job Description
Position Overview

You run the control plane where AIOps meets tenants - where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane - and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own
  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate
  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:
  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude - you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar Sr GPU Cloud K8S Expert (SRE SME) jobs: