Bachelor's degree in a relevant field such as Computer Science or Engineering.
5+ years of experience in Data Center operations or IT infrastructure management.
2+ years in team leadership or people management roles.
Experience managing 24x7 operations in mission-critical environments.
Strong knowledge of AI or HPC infrastructure, including NVIDIA technologies.
Responsibilities
Lead daily operations of the Data Center to ensure system reliability and excellence.
Manage a team of Operations Engineers, overseeing staffing and performance management.
Ensure 24x7 operational coverage to meet business needs.
Serve as the main escalation point for operational incidents and facilitate timely resolutions.
Oversee maintenance, troubleshooting, and operations of AI/HPC infrastructure components.
Establish and continuously improve Standard Operating Procedures and preventive maintenance programs.
Monitor site health and operational KPIs to drive service quality and efficiency.
Benefits
Opportunity to work with cutting-edge AI and HPC infrastructure.
Chance to lead a skilled team in a dynamic environment.
Support for professional development and continuous improvement initiatives.
Participation in a culture of operational excellence and teamwork.
Access to advanced technologies and systems in a mission-critical setting.
Full Job Description
Key Responsibilities
Lead and manage the daily operations of the Data Center site, ensuring the availability, reliability, and operational excellence of all infrastructure and systems.
Supervise and manage a team of Operations Engineers, including manpower planning, shift scheduling, task assignment, performance management, coaching, and professional development.
Ensure 24x7 operational coverage and maintain adequate staffing to support business and customer requirements.
Act as the primary escalation point for operational incidents and coordinate cross-functional teams to drive timely issue resolution and root cause analysis.
Oversee the operation, maintenance, and troubleshooting of AI/HPC infrastructure, including:
NVIDIA B300 Clusters
GPU Servers
x86 Servers
Storage Servers
Ethernet and InfiniBand Switches
Optical fiber, DAC, AOC, and associated cabling systems
Establish, maintain, and continuously improve operational procedures, Standard Operating Procedures (SOPs), Emergency Operating Procedures (EOPs), and preventive maintenance programs.
Monitor site health, operational KPIs, incident trends, and infrastructure performance to ensure service quality and operational efficiency.
Coordinate hardware installation, rack and stack activities, system commissioning, infrastructure expansion, and lifecycle management.
Review and approve maintenance activities, change requests, incident reports, and shift handover records.
Ensure compliance with company policies, operational standards, safety requirements, and security procedures within the Data Center.
Collaborate with engineering, network, facilities, and vendor teams to support new deployments and operational improvement initiatives.
Participate in on-call rotation and provide hands-on operational support when necessary, including covering shift duties during manpower shortages, emergencies, or critical incidents.
Drive a culture of operational excellence, teamwork, accountability, and continuous improvement within the site operation team.
Qualifications
Bachelor's degree or above in Computer Science, Computer Engineering, Electrical Engineering, Electronics Engineering, Information Technology, or related disciplines.
Minimum 5 years of experience in Data Center operations, IT infrastructure, or HPC/AI infrastructure management.
Minimum 2 years of experience in team leadership or people management.
Proven experience managing 24x7 shift operations in a mission-critical environment is preferred.
Experience with large-scale AI or HPC clusters is highly desirable.
Strong knowledge of Data Center operations and infrastructure management, including:
NVIDIA B300 Clusters
GPU Servers
x86 Servers
Storage Systems
Ethernet Networking
InfiniBand Networking
Familiarity with NVIDIA GPU architecture, NVLink, NVSwitch, and AI cluster deployment concepts.
Experience with server hardware troubleshooting, firmware management, and hardware lifecycle management.
Good understanding of structured cabling systems, including optical fiber, MPO/LC connectors, DAC, and AOC cabling.
Familiarity with infrastructure monitoring and management tools.
Linux and System Administration
Strong Linux administration and troubleshooting skills, including:
System and service management
Hardware and performance diagnostics
Log analysis and incident investigation
Network troubleshooting
Basic scripting and automation
Leadership and Management Skills
Demonstrated ability to lead, motivate, and develop a team of Operations Engineers.
Experience in:
Shift scheduling and workforce planning
Incident and escalation management
Performance management and coaching
SOP/EOP development and operational governance
Vendor and stakeholder coordination
Strong decision-making skills with the ability to manage priorities in a fast-paced and mission-critical environment.
Personal Attributes
Willingness to provide hands-on operational support and participate in on-call duties when required.
Strong ownership mindset and accountability for site operations.
Excellent communication and interpersonal skills.
Ability to remain calm and make sound decisions during critical incidents.
Highly organized, detail-oriented, and committed to operational excellence and continuous improvement.