Data Center Operations Engineer

Bitdeer Technologies Group

$80K — $95K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Engineering, or IT-related field.
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with NVIDIA B300 and various server types including GPU and x86 servers.
  • Knowledge of server hardware components and firmware management.
  • Basic Linux administration skills, especially in system monitoring and troubleshooting.

Responsibilities

  • Oversee daily operation and maintenance of the Data Center to ensure high service availability.
  • Install, rack, and troubleshoot AI/HPC infrastructure including various servers and switches.
  • Monitor the health status of all cluster systems and associated infrastructure.
  • Conduct hardware maintenance tasks including replacements and firmware upgrades.
  • Support server provisioning, operating system installations, and network validation.
  • Identify and resolve hardware and infrastructure issues promptly.
  • Maintain accurate logs of operations, maintenance activities, and incidents.

Benefits

  • Opportunity to work with cutting-edge AI Data Center technologies.
  • Exposure to NVIDIA B300 clusters and HPC infrastructure.
  • Collaboration with engineering and infrastructure teams for ongoing improvements.
  • Participation in regular training and professional development sessions.
Full Job Description
Key Responsibilities
  1. Responsible for the daily operation and maintenance of the Data Center infrastructure to ensure high availability and stable service operation.
  2. Perform installation, rack and stack, cabling, commissioning, maintenance, and troubleshooting of AI/HPC cluster infrastructure, including:
    • NVIDIA B300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Switches
    • DAC, AOC, Optical Fiber, and related cabling infrastructure
  3. Monitor and maintain the health status of cluster systems, including servers, GPUs, storage, networking devices, and associated infrastructure.
  4. Conduct hardware replacement and maintenance activities, including FRU replacement, BIOS/BMC/Firmware upgrades, and hardware diagnostics.
  5. Support server provisioning, operating system installation, cluster expansion, network validation, and burn-in testing.
  6. Troubleshoot hardware and infrastructure issues, including server failures, GPU errors, storage issues, network connectivity problems, switch failures, and cabling faults.
  7. Perform routine inspections, preventive maintenance, and maintain accurate operational records and maintenance logs.
  8. Execute incident response procedures and provide timely escalation and resolution according to operational standards.
  9. Prepare shift handover reports and maintain operation documents, SOPs, and incident reports.
  10. Work closely with engineering, network, and infrastructure teams to support new deployments and ongoing operation improvements.
  11. Participate in a three-shift rotation schedule, including night shifts, weekends, and holidays as required.

Requirements

Education
  • Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.

Technical Skills
  • Basic understanding of Data Center infrastructure and server hardware architecture.
  • Familiarity with one or more of the following systems:
    • NVIDIA B300 Cluster
    • GPU Servers
    • x86 Servers
    • Storage Servers
    • Ethernet and InfiniBand Networks
  • Knowledge of server hardware components, including CPU, memory, storage, GPU, BMC/IPMI, and firmware management.
  • Familiarity with network concepts, including TCP/IP, Ethernet, VLAN, Link Aggregation (LACP), and high-speed interconnect technologies such as InfiniBand or RoCE.
  • Understanding of structured cabling systems, including DAC, AOC, optical fiber, MPO, and LC connectors.

Linux Skills
  • Basic Linux administration skills, including:
    • System monitoring and troubleshooting
    • Service management using systemctl
    • Log analysis using journalctl and dmesg
    • Network troubleshooting tools such as ip and ethtool
    • Basic shell scripting

Preferred Qualifications
  • Experience in Data Center operations or hardware maintenance is preferred.
  • Experience supporting AI/HPC infrastructure or GPU clusters is a plus.
  • Familiarity with NVIDIA AI infrastructure, including GB200 and GB300 systems, is highly desirable.
  • Experience with large-scale cluster environments and high-speed networking technologies is a plus.
  • Familiarity with monitoring and orchestration tools such as Slurm, Kubernetes, Prometheus, or Grafana is an advantage.

Personal Attributes
  • Willingness to work in a 24x7 shift rotation schedule, including night shifts.
  • Strong sense of responsibility and ownership.
  • Good teamwork and communication skills.
  • Ability to work under pressure and respond effectively to operational incidents.
  • Detail-oriented with strong adherence to operational procedures and safety standards.
  • Self-motivated with a proactive attitude toward learning and problem-solving.

This position is ideal for candidates who are interested in building and operating next-generation AI Data Center infrastructure supporting large-scale NVIDIA B300 clusters.

Similar Jobs

More Jobs at Bitdeer Technologies Group

More Information Technology Jobs

Find similar Data Center Operations Engineer jobs: