Position OverviewYou run the control plane where AIOps meets tenants - where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.
NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane - and you make sure the AIOps substrate can reach in and remediate without a human on the pager.
What you'll own- Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs).
- Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
- Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
- AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
- Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
- Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
- Terraform providers and modules for infrastructure-as-code across GPU clusters.
- SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
- Incident management: runbook automation, escalation, post-incident reviews.
- Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
- GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.
Feed the AIOps substrate- The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
- Your CRDs are the schema the platform's predictors and remediators write against.
- Every human intervention you do this quarter becomes an autonomous workflow next quarter.
What success looks like in year 1- Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
- BMaaS live for external tenants with self-service onboarding.
- Cluster availability and job-completion SLOs published and met.
Job Requirement:- 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
- Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
- Experience with topology-aware scheduling and GPU-specific resource management
- Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
- Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
- Experience with Prometheus, Grafana, and alerting at scale
- Strong programming skills in Go or Python for operator/CRD development
- AIOps aptitude - you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
- Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.