You will build the high-performance compute environments we deliver to customers. These environments span bare-metal servers, virtual machines with GPU passthrough, Kubernetes and Slurm clusters, distributed storage, and high-speed networking.
You will work across the infrastructure stack - from Linux images and hardware provisioning to cluster networking, GPU performance, observability, and lifecycle automation. The goal is to make every customer environment performant, reliable, isolated, and repeatable.
Key Responsibilities:- Design, automate, validate, and deliver the complete lifecycle of customer compute environments - from provisioning through upgrades, recovery, and decommissioning
- Use AI aggressively to automate and accelerate every aspect of infrastructure delivery and operations
- Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads
- Build and maintain Linux images and automated OS-provisioning workflows
- Operate the NVIDIA GPU stack: drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring
- Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP
- Configure distributed and shared storage for high-performance workloads
- Build monitoring, alerting, diagnostics, and automated recovery for customer environments
- Develop reusable tooling, standards, documentation, and runbooks
- Collaborate with customers and internal teams to translate workload requirements into sound infrastructure designs
Requirements:- 5+ years of experience building and operating production Linux infrastructure
- Strong production experience with Kubernetes on bare metal (bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, troubleshooting)
- Experience with Linux virtualization: KVM/QEMU, libvirt, VFIO device passthrough
- Experience operating NVIDIA GPUs on Linux and Kubernetes (drivers, container runtimes, device plugins, GPU Operator, GPU telemetry)
- Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, packet-level troubleshooting (tcpdump, Wireshark)
- Practical scripting experience
- Experience with configuration-management tools such as Ansible
- Ability to diagnose complex, cross-layer infrastructure issues
- Strong communication and ability to drive technical decisions across teams
- Track record of moving quickly, taking ownership, and continuously improving systems
Nice to Have:- Production Slurm experience
- High-performance networking: NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX
- Hugepages, NUMA, CPU pinning
- SR-IOV, DPDK
- Distributed storage: Ceph, Lustre, Weka
- KubeVirt, OpenStack
- IPsec, WireGuard, Tailscale
- VXLAN, BGP, ECMP
- Bare-metal management: BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init
- Network automation: NetBox, Nautobot, Nornir
- AI training, inference, or distributed GPU workload infrastructure
- Python or Go proficiency