K8 Site Reliability SME

Bitdeer Technologies Group

$150K — $180K *
Technical Services
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S.
  • Deep understanding of Nvidia GPU operator and device plugins.
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees.
  • Experience with bare-metal server provisioning and lifecycle automation.
  • Proficiency in Terraform, Helm, and GitOps workflows.

Responsibilities

  • Design and deploy production Kubernetes clusters optimized for GPU workloads at scale.
  • Oversee topology-aware scheduling for GPU locality and NVLink domain awareness.
  • Manage AI framework integrations like Slurm and Kubeflow on K8S.
  • Automate incident management processes including runbook automation.
  • Implement monitoring solutions using Prometheus and Grafana.

Benefits

  • Flexible work environment with remote options.
  • Opportunity to work at the cutting edge of AI and GPU technology.
  • Professional development support for certifications and learning.
  • Strong emphasis on teamwork and collaboration.
  • Health and wellness programs to promote work-life balance.
Full Job Description
Position Overview

You run the control plane where AIOps meets tenants - where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane - and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own
  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate
  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:
  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude - you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Technical Services Jobs

Find similar K8 Site Reliability SME jobs: