Sr SRE & Automation Engineer (Customer Facing)

Bitdeer Technologies Group

$135K — $160K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in SRE/cloud operations with 2+ years managing GPU workloads at scale
  • Expertise in Kubernetes operations and GPU workload management
  • Hands-on experience with multi-tenant cloud platforms ensuring strong isolation
  • Strong proficiency in Terraform, Helm, and GitOps workflows
  • Experience in incident management with a focus on customer communications
  • Deep understanding of topology-aware scheduling for GPU resources
  • Programming skills in Go or Python for automation purposes

Responsibilities

  • Own end-to-end reliability for the GPU cloud service and tenant experience
  • Optimize production Kubernetes clusters for large-scale GPU workloads
  • Manage incident response and communication with customers during outages
  • Automate monitoring and observability with tools like Prometheus and Grafana
  • Drive improvements between customer issues and system-level changes
  • Implement automated provisioning for Bare-Metal-as-a-Service solutions
  • Develop self-service observability features for tenants to monitor their workloads

Benefits

  • Flexible work environment with opportunities for remote work
  • Exposure to cutting-edge AI and cloud technology
  • Chance to shape operational practices in a high-scale startup
  • Engagement with a dynamic team in a fast-paced environment
  • Opportunity for professional development and growth in the AI/Cloud space
Full Job Description
Job Description

NeoCloud is building an AI-operated GPU cloud - and because it is a customer-facing cloud service, reliability is the product. Tenants run mission-critical training, fine-tuning, and inference workloads on our GPU infrastructure and trust us with their SLAs. In this role you own the reliability of the customer-facing GPU cloud service end-to-end: from tenant onboarding and service provisioning, through workload execution, incident response, and post-incident recovery. You are the SRE who stands between raw infrastructure and the customer's experience - designing the observability, automation, and operational practices that make a 10,000-GPU cloud feel simple and dependable to the tenants who depend on it.

What You'll Own
  • End-to-end reliability of the customer-facing GPU cloud service - availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity - placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty - tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling - minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.

Customer-Facing Ownership
  • You are accountable for the customer's reliability experience - when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
  • Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
  • Build self-service observability that lets customers answer their own questions - status, quota, job health - reducing support load.

Feed the AIOps Substrate
  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter - turning customer-impacting incidents into self-healing events.

What Success Looks Like in Year 1
  • Customer-facing GPU cloud service SLAs published and met - availability, job completion, provisioning latency.
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
  • BMaaS live for external tenants with self-service onboarding.
  • MTTD and MTTR for customer-impacting incidents reduced through automation.
  • Tenant self-service observability live - customers can see their own job health, quota, and status.

Requirements
  • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
  • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
  • Experience with topology-aware scheduling and GPU-specific resource management.
  • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
  • Customer-facing cloud service experience - defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
  • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
  • Experience with Prometheus, Grafana, and alerting at scale.
  • Strong programming skills in Go or Python for automation / operator development.
  • AIOps aptitude - you view the control plane as an execution surface for automated remediation, not just a scheduler.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar Sr SRE & Automation Engineer (Customer Facing) jobs: