Role ImpactYou'll build the systems that turn bare-metal GPU servers into reliable, production-ready compute. Own the machine lifecycle from discovery and provisioning through validation, upgrades, repair, and secure reuse, reducing manual work as our fleet grows.
Core Technical Responsibilities- Build automated discovery, network boot, OS imaging, and configuration workflows for GPU servers
- Automate BIOS, BMC, NIC, GPU driver, and firmware configuration with staged rollouts and safe recovery paths
- Develop hardware inventory and lifecycle services that track machine identity, configuration, health, and readiness
- Create acceptance tests and burn-in workflows for GPUs, memory, storage, and interconnects before capacity enters production
- Integrate provisioning and health checks with SLURM, Kubernetes, and compute allocation systems
- Build observability, quarantine, repair, and re-provisioning workflows; partner with datacenter teams to resolve hardware failures
- Implement secure credential handling, tenant isolation, and data sanitization across the server lifecycle
Technical RequirementsRequired Experience- 3+ years of experience operating Linux servers or building bare-metal infrastructure automation in production
- Hands-on experience with PXE/iPXE, DHCP, image provisioning, and out-of-band management such as Redfish or IPMI
- Strong software engineering and debugging skills in Python, Go, or a comparable language, plus Bash
- Experience designing reliable automation that handles partial failures, retries, and configuration drift
- Ability to own operational incidents and collaborate across hardware, networking, and platform teams
Infrastructure Skills- Linux boot, systemd, kernel and driver troubleshooting, and OS image management
- Infrastructure automation with tools such as Ansible and Terraform
- GPU server diagnostics, PCIe topology, BMC telemetry, and firmware compatibility
- Network fundamentals including addressing, VLANs, DNS, and management networks
- Metrics, logs, alerting, and auditable configuration management
Nice to Have- Experience with large NVIDIA GPU fleets, DGX/HGX platforms, or heterogeneous server vendors
- Experience with MAAS, Ironic, Tinkerbell, or similar provisioning systems
- Kubernetes or SLURM node lifecycle integrations
- Hardware qualification, automated burn-in, and fleet health scoring
- Contributions to open-source infrastructure tooling
Growth OpportunityYou'll work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure. You'll collaborate with our world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.
We value expertise and customer obsession - if you're passionate about building reliable, high-performance GPU infrastructure and have a track record of successful large-scale deployments, we want to talk to you.
Apply now and join us in our mission to democratize access to planetary scale computing.
CompensationCash compensation range of $150,000-$300,000 plus equity incentives.