Job Summary:The primary purpose of the AI Server / Rack Architecture Engineer (focused on L10/L11 system-level execution) is to own the physical-to-logical system blueprint, hardware configuration rules, and deployment blueprints for integrated AI supercomputer racks. This role directly contributes by acting as the ultimate system architect on the floor. They translate the high-level, theoretical chip designs from silicon vendors and the hyper-specific, customized hardware demands of Tier-1 hyperscale's into a structured, validated, and executable rack integration matrix.
Duties and Responsibilities:- Define and optimize rack-level system architecture, integration topology, and technical blueprints for AI server platforms, ensuring compute, network, power, cooling, storage, and mechanical subsystems are properly integrated to meet performance, scalability, reliability, and deployment requirements.
- Identify and resolve design conflicts and interface issues across hardware, mechanical, thermal, power, network, and firmware subsystems by establishing clear technical boundaries, interface requirements, and integration standards to ensure overall system compatibility and stability.
- Translate hyperscale customer requirements into rack-level architecture and design specifications while evaluating alignment with applicable Open Compute Project (OCP) standards. Coordinate required design customizations and compliance activities to ensure solutions meet customer, platform, and deployment requirements.
- Analyze system telemetry, diagnostic data, and failure patterns to identify complex rack-level integration and structural issues. Lead technical triage and root-cause investigations with cross-functional engineering teams to determine corrective actions and improve system reliability.
- Manage architecture-related Engineering Change Orders (ECOs) and evaluate their impact on rack-level design, interfaces, performance, and integration. Provide structured feedback to R&D teams based on validation, failure analysis, and production findings to drive design improvements and prevent recurring issues.
Education:- Bachelor of Science (BS) Degree in Computer Engineering, Electrical Engineering, Systems Engineering, or Computer Science (paired with heavy, documented experience in computer hardware architecture or data center systems)
Experience:- 1+ years of High-Performance Computing (HPC) or Enterprise Server Architecture
- 1+ years of High-Speed Interconnect Fabric & Scale-Out Cluster Topologies
- 1+ years of Open Compute Project (OCP) & Data Center Deployment Compliance
Physical Demands & Working Conditions:- Frequently remains stationary while performing system architecture, design reviews, and data analysis, and occasionally moves throughout laboratory, testing, and office areas to support rack-level integration and validation activities.
- Occasionally operates or handles AI server/rack equipment, test instruments, components, tools, and related hardware during integration, troubleshooting, and validation activities.
- Occasionally bends, reaches, crouches, or positions self to access rack-level equipment and frequently observes and evaluates hardware configurations, system connections, and test conditions.
- Frequently exchanges technical information with R&D, hardware, mechanical, thermal, power, network, firmware, and other cross-functional teams to coordinate system integration and resolve technical issues.
- Frequently works in office, engineering laboratory, and server/test environments and may be exposed to operating equipment, elevated noise levels, temperature variations, and other equipment-related conditions; appropriate PPE may be required.