Job Summary:The AI Server/Rack System Engineer leads the end-to-end integration, configuration, and validation of high-density AI server and rack systems. This role serves as the technical bridge across compute nodes, high-speed networking fabrics such as NVLink and InfiniBand, storage, BIOS/BMC firmware, and liquid-cooling systems to ensure all subsystems operate together as a unified, stable, and production-ready system.
Duties and Responsibilities:- Lead the integration and configuration of compute nodes, networking, storage, firmware, power, and cooling subsystems into fully functional AI racks. Validate system-level connectivity, configuration, and interoperability to ensure the rack operates as a unified computing platform.
- Execute system diagnostics, stress testing, and burn-in activities to validate rack stability, performance, and reliability. Monitor system health, thermal parameters, error counters, and other key indicators to identify issues before deployment.
- Lead system-level troubleshooting and root-cause analysis for hardware, networking, firmware, thermal, and integration failures. Coordinate with electrical, mechanical, software, thermal, and R&D teams to drive timely resolution of cross-functional issues.
- Manage BIOS, BMC, firmware, and system configuration alignment across rack components. Verify configuration consistency, track changes, and support firmware bring-up and validation throughout the integration process.
- Document system configurations, test results, failures, corrective actions, and deployment requirements for engineering and manufacturing teams. Prepare technical escalation reports and support effective handover from engineering integration to production and deployment.
Education:- Bachelor of Science degree in Computer Engineering (CE), Electrical Engineering (EE), Systems Engineering, Computer Science (CS) with a strong hardware/networking background, or a related discipline is required.
Experience:- 2-4 years of experience with data center infrastructure, enterprise servers, AI server/rack systems, or L10/L11 integration. Strong understanding of AI rack architecture and system-level integration is required.
- 2-4 years of experience configuring and troubleshooting high-speed networking fabrics such as InfiniBand, NVLink, Ethernet, or similar technologies. Ability to diagnose connectivity, interoperability, and performance issues across interconnected systems.
- 1+ year of experience with server management architecture, BIOS/BMC, firmware bring-up, configuration management, and related management protocols. Ability to maintain consistent configurations across integrated rack systems.
- 2-3 years of experience performing hardware diagnostics, system stress testing, burn-in, and reliability validation. Strong ability to monitor rack health, thermal parameters, error counters, logs, and system utilities to identify abnormal conditions.
- Strong system-level troubleshooting and root-cause analysis skills, with the ability to coordinate across hardware, firmware, networking, thermal, mechanical, and software teams to resolve complex integration issues.