Overview
Velerity Compute builds and recertifies AI-capable GPU servers: DGX and HGX-class A100 and H100 systems, InfiniBand fabric, and the hyperscaler and enterprise decom gear that feeds both. Today, every technical question a customer or broker asks routes through people whose primary job is running the floor or closing the deal. Answers get slow, configs get quoted before anyone confirms they are buildable, and deals stall on questions that take an experienced tech ten minutes.
This role is the technical answer. You are the person Sales calls when a customer asks whether four HGX baseboards can be made into two sellable systems, what the fabric needs to look like, whether a rack of R750s can be reconfigured to spec, or whether the units on the floor will pass certification. You spend most of your week in front of Sales and customers, and enough of it in front of the hardware to keep your answers honest.
Responsibilities
Technical support to Sales (roughly 60 percent) - Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
- Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
- Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
- Own configuration and assembly quality for systems being built to a customer order
- Own burn-in and certification testing, and sign off on units as sellable
- Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor
Hands-on technical work (roughly 40 percent) - Perform BIOS, firmware, and BMC audits on inbound AI systems and hyperscaler gear
- Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
- Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
- Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
- Own configuration and assembly quality for systems being built to a customer order
- Own burn-in and certification testing, and sign off on units as sellable
- Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor
Making the answer repeatable - Document recurring technical questions into reusable reference material so Sales stops asking the same thing twice
- Maintain the internal config catalog: what Velerity builds, what it costs to build, and what it takes to build it
- Train sales staff on enough technical baseline to qualify opportunities without escalating everything
Qualifications
- 5 or more years in data center hardware engineering, field engineering, systems integration, or a hyperscaler or ODM environment
- Direct hands-on experience with NVIDIA AI systems: DGX or HGX A100 and H100 platforms, SXM and PCIe form factors, NVLink and NVSwitch topology
- Working knowledge of InfiniBand fabric: HDR and NDR, ConnectX adapters, Quantum switches, subnet management, cabling and transceiver rules
- Broad enterprise server fluency across the major OEM lines, not just accelerated platforms. You should be able to spec, configure, and troubleshoot across:
- Dell EMC PowerEdge, including the R and XE series, iDRAC, PERC controllers, and OpenManage
- HPE ProLiant and Apollo, including iLO, Smart Array, and Intelligent Provisioning
- Supermicro, including SuperServer and GPU chassis lines, IPMI, and the SuperMicro Update Manager
- Lenovo ThinkSystem, including XClarity
- Cisco UCS, including B and C series, fabric interconnects, and UCS Manager or Intersight service profiles
- ODM and hyperscaler gear: OCP form factors, Open Rack, and high-density sled and blade platforms from the major ODMs
- Able to read a mixed manifest of enterprise gear and say quickly what it is, what it is worth configuring into, and what it will take
- Firmware and low-level configuration: BIOS, BMC, IPMI, Redfish, and OEM-specific management tooling and firmware bundles
- GPU diagnostics and fault interpretation, including XID codes, ECC and row remap behavior, bus and link faults, and firmware level faults
- Data center power and cooling fundamentals: three-phase power, rack level power budgeting, PDU configuration, high-density air cooling, exposure to liquid cooling
- Demonstrated ability to explain technical constraints to a non-technical audience without either dumbing it down or burying them