Job DescriptionNeoCloud is building an AI-operated GPU cloud - and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery. You are the SRE who stands between raw infrastructure and the customer's experience - designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.
What You'll Own- End-to-end reliability of the customer-facing GPU cloud service - availability, job completion, provisioning latency, and tenant experience.
- Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs) as the runtime substrate for customer workloads.
- Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
- Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity - placing customer jobs on the right hardware.
- Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
- Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
- SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
- Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
- Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty - tenant-aware dashboards and alerting.
- GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling - minimizing customer-visible impact.
- Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
- Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.
Customer-Facing Ownership- You are accountable for the customer's reliability experience - when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
- Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
- Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
- Build self-service observability that lets customers answer their own questions - status, quota, job health - reducing support load.
Feed the AIOps Substrate- The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
- Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
- Every human intervention you do this quarter becomes an autonomous workflow next quarter - turning customer-impacting incidents into self-healing events.
What Success Looks Like in Year 1- Customer-facing GPU cloud service SLAs published and met - availability, job completion, provisioning latency.
- Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
- BMaaS live for external tenants with self-service onboarding.
- MTTD and MTTR for customer-impacting incidents reduced through automation.
- Tenant self-service observability live - customers can see their own job health, quota, and status.
Requirements- 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
- Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
- Experience with topology-aware scheduling and GPU-specific resource management.
- Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
- Customer-facing cloud service experience - defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
- Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
- Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
- Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
- Experience with Prometheus, Grafana, and alerting at scale.
- Strong programming skills in Go or Python for automation / operator development.
- AIOps aptitude - you view the control plane as an execution surface for automated remediation, not just a scheduler.
- Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.