Server Engineer - AI Systems

Sprout

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years in data center hardware engineering or systems integration
  • Hands-on experience with NVIDIA AI systems (DGX/HGX A100 and H100)
  • Knowledge of InfiniBand fabric and related adapters and switches
  • Fluency with major enterprise server lines (Dell EMC, HPE, Supermicro, Lenovo, Cisco)
  • Ability to diagnose and interpret GPU faults and field diagnostics
  • Understanding of data center power and cooling fundamentals
  • Strong communication skills for technical explanations to non-technical audiences

Responsibilities

  • Provide technical support to Sales for customer inquiries
  • Run GPU diagnostics and troubleshoot hardware faults
  • Configure and validate InfiniBand and Ethernet fabric
  • Ensure quality in assembly for build-to-order systems
  • Conduct burn-in and certification testing of units
  • Document and maintain technical knowledge for sales staff
  • Train Sales staff on technical qualifications to reduce escalations

Benefits

  • Collaborative team environment with top-tier professionals
  • Opportunity for hands-on work with cutting-edge AI systems
  • Direct impact on customer satisfaction and sales success
  • Access to advanced tools and resources in the data center space
  • Training and development opportunities to enhance technical skills
Full Job Description
Overview

Velerity Compute builds and recertifies AI-capable GPU servers: DGX and HGX-class A100 and H100 systems, InfiniBand fabric, and the hyperscaler and enterprise decom gear that feeds both. Today, every technical question a customer or broker asks routes through people whose primary job is running the floor or closing the deal. Answers get slow, configs get quoted before anyone confirms they are buildable, and deals stall on questions that take an experienced tech ten minutes.

This role is the technical answer. You are the person Sales calls when a customer asks whether four HGX baseboards can be made into two sellable systems, what the fabric needs to look like, whether a rack of R750s can be reconfigured to spec, or whether the units on the floor will pass certification. You spend most of your week in front of Sales and customers, and enough of it in front of the hardware to keep your answers honest.

Responsibilities

Technical support to Sales (roughly 60 percent)
  • Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
  • Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
  • Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
  • Own configuration and assembly quality for systems being built to a customer order
  • Own burn-in and certification testing, and sign off on units as sellable
  • Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor

Hands-on technical work (roughly 40 percent)
  • Perform BIOS, firmware, and BMC audits on inbound AI systems and hyperscaler gear
  • Run GPU diagnostics and interpret results, including NVIDIA field diagnostics, DCGM, and XID fault classification, and make disposition calls on that basis
  • Troubleshoot faults across GPU, baseboard, NVLink and NVSwitch, PSU, cooling, and fabric
  • Configure and validate InfiniBand and Ethernet fabric, including adapter firmware, cabling, and transceiver compatibility
  • Own configuration and assembly quality for systems being built to a customer order
  • Own burn-in and certification testing, and sign off on units as sellable
  • Feed accurate technical attributes back into the master data and inventory record so what Sales sees is what is on the floor

Making the answer repeatable
  • Document recurring technical questions into reusable reference material so Sales stops asking the same thing twice
  • Maintain the internal config catalog: what Velerity builds, what it costs to build, and what it takes to build it
  • Train sales staff on enough technical baseline to qualify opportunities without escalating everything


Qualifications

  • 5 or more years in data center hardware engineering, field engineering, systems integration, or a hyperscaler or ODM environment
  • Direct hands-on experience with NVIDIA AI systems: DGX or HGX A100 and H100 platforms, SXM and PCIe form factors, NVLink and NVSwitch topology
  • Working knowledge of InfiniBand fabric: HDR and NDR, ConnectX adapters, Quantum switches, subnet management, cabling and transceiver rules
  • Broad enterprise server fluency across the major OEM lines, not just accelerated platforms. You should be able to spec, configure, and troubleshoot across:
  • Dell EMC PowerEdge, including the R and XE series, iDRAC, PERC controllers, and OpenManage
  • HPE ProLiant and Apollo, including iLO, Smart Array, and Intelligent Provisioning
  • Supermicro, including SuperServer and GPU chassis lines, IPMI, and the SuperMicro Update Manager
  • Lenovo ThinkSystem, including XClarity
  • Cisco UCS, including B and C series, fabric interconnects, and UCS Manager or Intersight service profiles
  • ODM and hyperscaler gear: OCP form factors, Open Rack, and high-density sled and blade platforms from the major ODMs
  • Able to read a mixed manifest of enterprise gear and say quickly what it is, what it is worth configuring into, and what it will take
  • Firmware and low-level configuration: BIOS, BMC, IPMI, Redfish, and OEM-specific management tooling and firmware bundles
  • GPU diagnostics and fault interpretation, including XID codes, ECC and row remap behavior, bus and link faults, and firmware level faults
  • Data center power and cooling fundamentals: three-phase power, rack level power budgeting, PDU configuration, high-density air cooling, exposure to liquid cooling
  • Demonstrated ability to explain technical constraints to a non-technical audience without either dumbing it down or burying them


Similar Jobs

More Jobs at Sprout

  • Operations Manager
    $80K — $95K *
    Garland, TX 75040 (Dallas County)
    Manufacturing & Automotive
    In-Person
  • AI Server Configuration Lead
    $110K — $130K *
    Garland, TX 75040 (Dallas County)
    Information Technology
    In-Person
  • AI Enablement Lead
    $110K — $130K *
    Garland, TX 75040 (Dallas County)
    Enterprise Technology
    In-Person
  • Process Engineer
    $80K — $95K *
    Garland, TX 75040 (Dallas County)
    Manufacturing & Automotive
    In-Person
  • Business Systems Analyst
    $90K — $110K *
    Garland, TX 75040 (Dallas County)
    Business Services
    In-Person

More Information Technology Jobs

Find similar Server Engineer - AI Systems jobs: