About the Role:We are seeking a highly skilled and motivated GPU Fleet Operations Engineer to join Crusoe's Fleet Operations team. This role is focused on the advanced diagnosis, maintenance, and repair of high-performance GPU compute clusters, ensuring maximum uptime, reliability, and performance across our fleet.
The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs. This position plays a critical role in maintaining the health and scalability of Crusoe's rapidly growing GPU fleet.
What You'll Be Doing:- Automate deep-level diagnosis and troubleshooting of hardware faults within GPU racks and high-density compute systems.
- Develop software to troubleshoot and support GPU platforms including NVIDIA A100, H200, GB200, B200, B300 and AMD 350X / 355X.
- Execute component-level diagnosis and remediation for failed or degraded hardware.
- Partner with data center operations to manage and perform field-replaceable unit (FRU) repairs for GPUs, power supplies, cooling systems, interconnects, and networking hardware.
- Conduct post-repair validation, burn-in testing, torch testing, and NVIDIA NCCL testing to ensure system stability and performance.
- Implement and execute preventative maintenance procedures to improve fleet reliability and extend hardware lifespan.
- Perform firmware and BIOS upgrades across the GPU fleet.
- Maintain detailed documentation of maintenance activities, failures, and resolutions in ticketing and asset management systems.
- Develop and update standard operating procedures (SOPs) for troubleshooting, repair, and validation workflows.
- Collaborate with engineering, software, and data center operations teams to identify root causes of systemic failures and implement preventative solutions.
- Participate in a rotating infrastructure on-call schedule (about one week every 4-6 weeks) with daytime coverage and handoff to the Europe team.
What You'll Bring to the Team:- Ability to code in Golang
- Proven experience diagnosing and repairing high-density, rack-mounted compute hardware in production environments.
- Deep understanding of GPU architectures and hands-on experience with GPU-based systems.
- Experience supporting NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X series platforms.
- Familiarity with high-speed interconnects such as InfiniBand, NVLink, and RDMA over Converged Ethernet (RoCE).
- Strong Linux experience (Ubuntu, Rocky Linux, CentOS) using the command line for diagnostics and testing.
- Proficiency with GPU and system diagnostic tools such as NVIDIA DCGM and NVIDIA field diagnostic utilities.
- Experience working with enterprise server hardware, power delivery, and cooling systems.
- Strong analytical and problem-solving skills.
- Excellent communication and collaboration skills.
- Ability to work independently in a fast-paced data center or operations environment.
Nice to Have:- Technical certification or Associate's/Bachelor's degree in Electrical Engineering, Computer Science, or a related field or demonstrated experience.
- Experience working directly with hardware vendors and escalations.
- Background in large-scale GPU fleet operations or hyperscale data center environments.
Benefits:- Hybrid work schedule
- Industry competitive pay
- Restricted Stock Units in a fast growing, well-funded technology company
- Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
- Employer contributions to HSA accounts
- Paid Parental Leave
- Paid life insurance, short-term and long-term disability
- Teladoc
- 401(k) with a 100% match up to 4% of salary
- Generous paid time off and holiday schedule
- Cell phone reimbursement
- Tuition reimbursement
- Subscription to the Calm app
- MetLife Legal
- Company paid commuter benefit; $300 per pay period
Compensation Range:Compensation will be paid in the range of $215,000 - $260,000. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.