Celestica

Principal Hardware Systems Architect-Compute Platforms

Celestica$162K — $195K *
Telecommunications & Hardware
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's in a related technical field (Electrical Engineering, Computer Engineering, Computer Science, Physics)
  • 12+ years of experience in hardware and system development (focus on server, AI, HPC, or data-center platforms)
  • Proven experience leading architecture for complex compute platforms
  • Strong knowledge of x86/ARM processors, GPUs, PCIe, CXL, and high-speed interconnects
  • Expertise in the relationships between compute, networking, memory, power, and thermal architecture
  • Experience in multi-disciplinary collaboration across engineering teams
  • Excellent communication and presentation skills

Responsibilities

  • Lead the architecture development for next-generation AI and high-performance computing platforms
  • Define and own architectural frameworks for CPUs, GPUs, memory, and storage subsystems
  • Establish system requirements and evaluate architectural trade-offs in performance and cost
  • Maintain an up-to-date understanding of evolving compute and AI technologies
  • Collaborate with peer architects to define integrated architectures across systems
  • Drive architectural strategies from compute nodes to multi-rack systems
  • Partner with engineering teams to establish comprehensive system requirements and interfaces

Benefits

  • Collaborative and innovative team environment
  • Opportunity to work on next-gen AI and high-performance computing technology
  • Exposure to cutting-edge technologies and leading compute ecosystem suppliers
  • Potential for influencing major architectural decisions and working closely with industry leaders
  • Focus on professional development and cross-disciplinary learning
Full Job Description
Req ID: 140082
Region: Americas
Country: USA
State/Province: California
City: San Jose

Summary

Celestica is seeking a Principal Hardware Systems Architect for Compute Platforms to provide senior technical leadership for the architecture and development of next-generation AI, accelerated computing, and high-performance data center platforms.

This role will lead compute-platform architecture spanning CPUs, GPUs and AI accelerators, memory, storage, PCIe/CXL, scale-up interconnects, and compute-node/tray architecture, while collaborating with peer domain architects and engineering leaders to define complete rack and multi-rack systems incorporating networking, power, thermal, mechanical, firmware, and system-management technologies.

This is a system and platform architecture role, not a processor or ASIC design role. The successful candidate will operate above the component level, translating customer requirements, workload needs, and technology roadmaps into scalable compute architectures.

A critical aspect of the position is understanding that next-generation AI infrastructure can no longer be optimized as independent compute, networking, power, and cooling subsystems. The Compute Architect will partner closely with Celestica's Network Architect and other domain experts to optimize the complete AI system from compute node through rack and multi-rack deployments.

Skills & Responsibilities:

Compute Platform Architecture
  • Define architecture for next-generation AI, accelerated-compute, server, and high-performance computing platforms.
  • Own architectural definition of:
    • CPU architectures
    • GPU and AI accelerator subsystems
    • Memory and HBM architectures
    • PCIe and CXL fabrics
    • Local storage and NVMe architectures
    • GPU/accelerator topology
    • Scale-up interconnect architecture
    • Compute node and tray architecture
  • Establish compute-system requirements, interfaces, partitioning, and architectural tradeoffs.
  • Evaluate alternatives based on performance, bandwidth, latency, power, thermal density, reliability, cost, scalability, serviceability, and manufacturability.
  • Develop reusable compute-platform architectures and building blocks that can be leveraged across customers and programs.


AI & Compute Technology Leadership
  • Maintain deep technical understanding of evolving CPU, GPU, AI accelerator, memory, storage, and interconnect technologies.
  • Track and evaluate technology roadmaps from leading compute ecosystem suppliers including NVIDIA, AMD, Intel, ARM ecosystem companies, and emerging AI accelerator providers.
  • Understand emerging AI compute architectures and their implications for:
    • Accelerator density
    • Memory bandwidth
    • PCIe/CXL connectivity
    • Scale-up fabrics
    • Scale-out networking
    • Rack power
    • Liquid cooling
    • Rack and cluster architecture
  • Lead architectural evaluation of scale-up technologies, including NVLink/NVSwitch-class technologies and emerging accelerator interconnects.


Integrated Rack & Multi-Rack Architecture
  • Partner with peer system architects to define integrated rack and multi-rack architectures spanning compute, networking, power, thermal, mechanical, firmware, and system-management domains.
  • Optimize architecture as a complete AI system rather than as independent compute or networking platforms.
  • Partner closely with the Network Architect to establish the interface between accelerator scale-up architectures and network scale-out architectures.
  • Jointly evaluate system-level architectural tradeoffs where compute and networking intersect, including:
    • Accelerator density
    • Fabric bandwidth
    • NIC/DPU architecture
    • Switch topology
    • Latency
    • Power consumption
    • Thermal density
    • Reliability
    • Cost
    • Scalability
  • Work with Power, Thermal, Mechanical, Reliability, Firmware, and Validation engineering teams to establish system-level requirements and interfaces.
  • Drive architectural thinking from compute node  tray  rack  multi-rack AI system.


Experience Requirements

Required Qualifications
  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, Physics, or related technical field.
  • Typically 12+ years of relevant hardware and system-development experience, with significant experience in server, compute, AI, HPC, or data-center platforms.
  • Demonstrated experience defining or leading system-level architecture for complex compute platforms.
  • Strong knowledge of modern server architectures including x86 and/or ARM processors, GPUs/accelerators, DDR/HBM, PCIe, CXL, NVMe/storage, and high-speed interconnects.
  • Strong understanding of the relationships among compute, networking, memory, power, and thermal architecture.
  • Experience taking complex platforms from early architecture through development, validation, and production.
  • Ability to make and communicate system-level engineering tradeoffs involving performance, cost, power, thermal, reliability, and schedule.
  • Strong technical communication and presentation skills.
  • Ability to influence technical decisions across multiple engineering disciplines without relying on direct organizational authority.
  • Experience engaging directly with customers, silicon vendors, and technology partners.


Preferred Qualifications
  • Experience architecting GPU-based AI training or inference platforms.
  • Experience with NVIDIA, AMD, Intel, or other accelerator-based platforms.
  • Experience with rack-scale AI systems and high-density compute architectures.
  • Knowledge of NVLink/NVSwitch-class technologies or comparable scale-up fabrics.
  • Working knowledge of high-performance Ethernet and/or InfiniBand scale-out architectures.
  • Experience with advanced liquid-cooled compute platforms.
  • Experience with high-power rack architectures reaching hundreds of kilowatts per rack.
  • Experience working with hyperscale or cloud-service-provider customers.
  • Experience developing reusable platform architectures.
  • Experience contributing to technology roadmaps, invention disclosures, patents, or advanced-development programs.


Typical Education

  • Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, Physics, or related technical field.
  • Educational Requirements may vary by Geography


Notes

This job description is not intended to be an exhaustive list of all duties and responsibilities of the position. Employees are held accountable for all duties of the job. Job duties and the % of time identified for any function are subject to change at any time.

About Celestica

Celestica is a Canadian multinational electronics manufacturing services company headquartered in Toronto, Ontario. The company provides a range of services to original equipment manufacturers (OEMs) in the aerospace and defense, communications, enterprise computing, healthcare, industrial, semiconductor, and smart energy industries. Celestica's services include design and engineering, supply chain management, assembly and testing, and after-market services. The company operates in North America, Europe, and Asia and has manufacturing facilities in over 10 countries. Celestica was founded in 1994 as a subsidiary of IBM Canada and became an independent company in 1997.
Learn more about Celestica
Size
23,915 employees
Market Cap
$1.3 billion
Industry
Founded
1994
5 Year Trend
-1.3%
NASDAQ

Similar Jobs

More Jobs at Celestica

More Telecommunications & Hardware Jobs

Find similar Principal Hardware Systems Architect-Compute Platforms jobs: