Job Summary:The AI Server/Rack System Lead Manager leads a multidisciplinary engineering team across System, Electrical, Power, and Mechanical Engineering to ensure multi-million-dollar AI hardware clusters are transformed from integrated racks into fully configured, stable, high-quality systems ready for deployment at hyperscale data centers. This role provides technical leadership across system integration, manufacturing execution, quality, customer requirements, and production readiness.
Duties and Responsibilities:- Lead System, Electrical, Power, and Mechanical Engineering teams in resolving complex production and system-level issues. Serve as the primary technical escalation point for critical issues on the manufacturing floor and drive timely resolution.Lead System, Electrical, Power, and Mechanical Engineering teams in resolving complex production and system-level issues. Serve as the primary technical escalation point for critical issues on the manufacturing floor and drive timely resolution.
- Drive improvements in factory yield, cycle time, system quality, and production efficiency for AI server/rack systems. Establish engineering quality standards and use data-driven analysis to identify and eliminate recurring production issues.
- Work directly with hyperscale's and key customers to align system requirements, validation criteria, and deployment expectations. Lead technical readiness reviews and provide engineering support for customer release and system sign-off.
- Manage engineering changes and ensure manufacturing feedback is incorporated into system and product improvements. Establish effective feedback loops between R&D, manufacturing, quality, and customer teams to improve product performance and reliability.
- Ensure engineering and manufacturing activities comply with high-voltage, mechanical, and workplace safety requirements. Promote safe engineering practices and enforce applicable U.S. regulatory and company safety standards.
Education:- Bachelor of Science (BS) degree in Electrical Engineering, Computer Engineering, Systems Engineering, Mechanical Engineering, Manufacturing Engineering, or a closely related field is required.
Experience:- 4-7+ years of experience in enterprise server, rack infrastructure, or AI cluster system engineering, with strong knowledge of full-stack AI cluster architecture, high-speed network topologies, and advanced infrastructure limitations.
- 3-5 years of experience leading multidisciplinary engineering teams, managing personnel, allocating resources, and resolving complex technical issues in fast-paced production environments.
- 2+ years of experience integrating high-density power and liquid cooling infrastructure, with the ability to understand system-level electrical, thermal, mechanical, and performance requirement2+ years of experience integrating high-density power and liquid cooling infrastructure, with the ability to understand system-level electrical, thermal, mechanical, and performance requirement.
- 3+ years of experience in manufacturing operations, yield analysis, cycle-time optimization, and quality frameworks, with demonstrated ability to drive root-cause analysis and continuous improvement.
- 2+ years of experience enforcing U.S. industrial safety requirements, with knowledge of applicable regulatory standards, high-voltage safety, mechanical safety, and the ability to lead technical decisions under pressure.