SRE L1 Support/Cloud Platform Ops Engineer

Bitdeer Technologies Group

$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration knowledge
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios)
  • Experience with ticketing systems (ServiceNow, Jira)
  • Ability to perform physical tasks in a data center (rack and stack, cabling)
  • Strong communication skills for effective handoffs and documentation
  • Curiosity about automation and process improvement

Responsibilities

  • Monitor GPU cluster health and network status via dashboards
  • Respond to alerts and execute runbooks for incidents
  • Perform hardware triage to identify failures
  • Execute standard remediation procedures on hardware
  • Manage incident tickets through resolution or escalation
  • Collect diagnostic data for escalations to L2/SME
  • Perform physical data center tasks and assist with hardware deployment

Benefits

  • Structured growth path leading to SME or automation roles
  • Opportunity to impact the evolution of AIOps automation
  • Hands-on experience with cutting-edge GPU technology
  • Engaging role that combines monitoring with process improvement
  • Collaborative environment with regular interaction with the APAC operations team
Full Job Description
Position Overview

You are the first human in the loop - the escalation target when the AIOps system needs a decision, and the source of ground truth that turns novel incidents into new automations.

NeoCloud is building an AI-operated GPU cloud. That doesn't mean fewer humans - it means humans focus on judgment calls the platform can't yet make, and every judgment call trains the platform to do it next time. In this L1 role you cover front-line monitoring and incident response for NeoCloud's US GPU DCs during the 8AM-8PM PST shift. You execute SOPs, escalate the hard cases, and feed the AIOps substrate the ground truth it needs to learn from novel incidents.

What you'll own
  • Monitor GPU cluster health, network status, storage systems, and environmental sensors via centralized dashboards.
  • Respond to alerts and execute runbooks for common incidents: GPU errors, link flaps, node failures, storage alerts.
  • Perform hardware triage: identify failed GPUs, NICs, PSUs, disks, and cables from monitoring data and physical inspection.
  • Execute standard remediation: GPU reset, node drain/reboot, link re-seat, BMC recovery.
  • Collect diagnostic data for L2/SME escalation: logs, DCGM output, network diagnostics, hardware health reports.
  • Manage incident tickets from creation through resolution or escalation (ServiceNow/Jira).
  • Perform physical DC tasks: cable installation, hardware swap-outs, rack and stack, labeling (on-site roles).
  • Execute structured shift handoffs at 8AM and 8PM PST with the APAC operations team.
  • Maintain and update operational runbooks based on recurring issues.
  • Assist with hardware deployment, firmware updates, and inventory management under SME guidance.

Feed the AIOps substrate
  • Every novel incident you resolve is data the platform team needs - you tag it, describe it, and hand it back so it becomes an automation.
  • Every runbook you touch should get closer to being executable by the platform, not by you.
  • Your handoff notes are structured signal, not free-form email.

Why this role is different from a NOC job
  • You are not the last line of defense - the platform is. You are the training signal.
  • Growth path is real: strong L1s here move into SME roles, or into the platform team as automation authors.

Job Requirement:
  • 2+ years in NOC, data center operations, or IT support role
  • Basic Linux system administration (command line, log analysis, service management)
  • Familiarity with monitoring tools (Prometheus, Grafana, Nagios, or equivalent)
  • Experience with ticketing systems (ServiceNow, Jira Service Management)
  • Ability to perform physical data center tasks: rack and stack, cabling, hardware replacement
  • Strong communication skills for shift handoffs, incident documentation, and escalation
  • Ability to work 8AM-8PM PST shift schedule (12-hour shifts with rotation)
  • Curiosity about automation - you don't just execute the runbook, you notice when it's the third time this month and ask what should change.
  • Comfort with structured data - you understand that how you file a ticket matters, because it may train a model that decides how the next one is filed.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar SRE L1 Support/Cloud Platform Ops Engineer jobs: