Minimum Qualifications- 10+ years in HPC, AI, or data center infrastructure engineering, including 7+ years in a customer-facing architecture role.
- Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Deep working knowledge of NVIDIA GPU platforms across the Hopper and Blackwell architectures, including Grace Hopper and Grace Blackwell superchip systems, and the associated software stack (CUDA, cuDNN, NCCL).
- Strong Linux systems engineering skills, including distributions tuned for HPC and AI workloads (Ubuntu, RHEL), kernel tuning, driver management, and system hardening.
- Automation and infrastructure-as-code experience, with proficiency in Ansible and Python, and familiarity with declarative tooling such as Terraform.
- Demonstrated ability to translate technical architecture into commercial outcomes for senior stakeholders.
Preferred Qualifications - Industry background: experience within a systems integrator (SI) or managed service provider (MSP) environment.
- Cluster management: hands-on experience with NVIDIA Base Command Manager and/or NVIDIA Mission Control.
- Network integration: understanding of high-speed interconnects (InfiniBand NDR/HDR, RoCEv2) and how they interact with host PCIe and NVLink topologies.
- Container platforms: Kubernetes and Red Hat OpenShift installation and administration.
- Cloud architecture: AI-oriented compute design on AWS, Azure, Google Cloud, or OCI.
- AI and orchestration tooling: exposure to Rafay, NVIDIA Run:ai, NVIDIA Omniverse, Red Hat OpenShift AI, and Kubeflow.
Certain states and localities require employers to post a reasonable estimate of salary range. A reasonable estimate of the current base pay range for this position is $110,000.00 to $150,000.00 annually. Actual salary will be based on a variety of factors, including shift, location, experience, skill set, performance, licensure and certification, and business needs. The range for this position in other geographic locations may differ. Certain positions may also be eligible for variable incentive compensation, such as bonuses or commissions, that is not included in the base pay.
The well-being of WWT employees is essential. When it comes to our benefits package, WWT has one of the best. We offer the following benefits to all full-time employees:
- Health and Wellbeing: Health (Medical & Prescription), Dental, and Vision Care, Onsite Health Centers (MO & IL), Employee Assistance Program, Wellness program
- Financial Benefits: Competitive Pay, Profit Sharing, 401k Plan with Company Matching, Life and Disability Insurance, Flexible Spending Accounts, Tuition Reimbursement
- Paid Time Off: PTO & Holidays, Parental Leave, Medical Leave, Military Leave, Bereavement, Day of Caring
- Additional Perks: Family Planning Benefits, Nursing Mothers Benefits, Voluntary Legal, Voluntary Supplemental Accident/Illness/Hospital, Voluntary ID Theft, Pet Insurance, Employee Discount Program
Note: This is not an all-encompassing list and should not be used as a complete description of the plan's benefits. For more information, see our US benefits website at wwt.com/us-benefits.
About the Role As Domain Architect - AI Compute, you will be the primary technical authority for the physical and logical lifecycle of high-performance GPU compute fleets across a diverse portfolio of client environments, bridging the gap between architectural design and hands-on execution. You are a builder as much as an advisor: as comfortable configuring a cluster from the CLI as you are explaining that configuration to a C-level audience.
As a global systems integrator, we don't simply operate static cloud environments. We design and deliver purpose-built, high-scale AI factories for some of the world's leading enterprises. In this role you will define the reference standard for compute infrastructure, moving beyond single-server administration to architect repeatable, scalable, and automated compute fabrics. You will act as technical lead on NVIDIA Cloud Partner (NCP) and private enterprise AI cloud deployments, owning the compute layer of the compute, network, and storage stack.
Your time will be split roughly 60/40 between delivering complex AI infrastructure (60%) and providing pre-sales subject matter expertise (40%). You will lead the physical provisioning of NVIDIA DGX SuperPOD, NVIDIA DGX BasePOD, and Cisco AI POD environments, ensuring clients inherit platforms that are genuinely ready for day-2 operations, while helping the sales team scope and cost future deployments.
Key Responsibilities Delivery and implementation Bare-metal build and provisioning - Lead the physical provisioning of GPU platforms and clusters, including NVIDIA GB200/GB300 NVL72 rack-scale systems, NVIDIA DGX SuperPOD and DGX BasePOD reference architectures, HGX- and MGX-based systems, and Cisco AI PODs.
- Use NVIDIA Base Command Manager (BCM) for cluster provisioning, diskless boot, image management, and firmware lifecycle management.
- Harden and baseline the host operating system across the fleet.
- Establish cluster monitoring and lifecycle management with NVIDIA Mission Control.
- Build and execute zero-touch provisioning (ZTP) workflows that turn bare-metal hardware into production-ready nodes.
Scheduler and workload configuration - Define and enforce fair-share policies, fractional GPU allocation using Multi-Instance GPU (MIG), and preemption logic for multi-tenant environments.
- Ensure orchestration layers respect hardware topology, including NUMA affinity and PCIe topology, to protect performance.
- Implement and tune advanced schedulers: Slurm on bare metal, and NVIDIA Run:ai, Kueue, or Volcano on Kubernetes.
Orchestration and day-2 operations - Deploy and configure multi-cluster management and observability platforms such as Rafay.
- Implement high-fidelity telemetry using NVIDIA Data Center GPU Manager (DCGM) to monitor GPU health, thermal throttling, and Xid error rates.
- Lead the transition to day-2 operations, ensuring the environment is fully integrated with the client's identity providers and storage backends before handoff.
Performance engineering - Conduct acceptance and validation testing using NCCL tests and HPL to verify cluster performance against expected baselines.
- Perform kernel and OS-level tuning (hugepages, sysctl, and driver parameters) to optimize for high-bandwidth InfiniBand and RoCEv2 fabrics.
Pre-sales SME and consulting Technical scoping and estimation - Support the sales team by validating customer technical requirements and producing accurate level-of-effort (LOE) estimates for statements of work.
- Define the standard operating environment (SOE) used in proposals to ensure repeatability across engagements.
Bill of materials validation and architecture - Own the technical accuracy of the compute bill of materials.
- Verify that all components, including memory, NVMe storage, NICs, and optical transceivers, are aligned with NVIDIA-Certified Systems requirements and the qualified component lists in the relevant DGX, HGX, MGX, and NVL72 reference architectures.
- Match GPU platform selection to the client's actual workload profile and growth expectations.
Client workshops - Lead technical discovery workshops to determine workload requirements, distinguishing between the very different infrastructure profiles of traditional machine learning, LLM training, LLM inference, and Omniverse digital twin and visualization workloads.