Senior Software Engineer (DCIE)

Crusoe

$170K — $205K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4-6 years of software engineering experience
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP)
  • Strength in at least one programming language - Go, Python, Java, Rust
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration skills
  • Ability to work independently and in a team setting

Responsibilities

  • Develop and implement deep-level diagnostics and troubleshooting of hardware faults within GPU racks
  • Create troubleshooting and automation tooling for GPU platforms like NVIDIA and AMD
  • Develop automation and AI agents for diagnosing and remediating hardware failures
  • Collaborate with data center operations to innovate tooling and AI agents for environment management
  • Develop validation and testing tools for system stability and performance post-repair
  • Oversee the deployment, monitoring, and operational support of developed tooling
  • Create automation and operational tooling for facilities management and cooling systems

Benefits

  • Industry competitive pay
  • Restricted Stock Units in a fast-growing technology company
  • Health insurance options including HDHP, PPO, vision, and dental for you and dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short-term and long-term disability
  • Teladoc access
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal services
  • Company paid commuter benefit of $50 per pay period
Full Job Description
About the Role:

We are seeking a highly skilled and motivated Software Engineer to join Crusoe's Data Center Infrastructure Engineering team. This position is focused on the development of software for the management of a fleet of GPU servers as well as the data centers that house those systems. The role focuses on the developing and implementing advanced diagnostic, observability, automation and repair tooling for high-performance GPU compute clusters.

The ideal new team member will be a hands-on problem solver who is comfortable working independently. The new team member will play a critical role in maintaining the health and scalability of Crusoe's rapidly growing GPU fleet.

What You'll Be Doing:
  • Developing and implementing deep-level diagnostics and troubleshooting of hardware faults within GPU racks and high-density compute systems.
  • Developing troubleshooting and automation tooling for GPU platforms including NVIDIA A100, H200, GB200, B200 and AMD 350X / 355X.
  • Developing automation and AI agents for executing component-level diagnosis and remediation for failed or degraded hardware.
  • In conjunction with data center operations develop innovative tooling and AI agents for managing the critical environment.
  • Developing tooling for post-repair validation and testing tools such as burn-in, Pytorch, and NVIDIA NCCL to ensure system stability and performance.
  • Own the deployment, monitoring, and operational support of developed tooling, ensuring solutions maximize GPU fleet availability and performance to drive customer success.
  • Developing automation and operational tooling for facilities management power as well as direct liquid cooling hardware systems


What You'll Bring to the Team:
  • 4-6 years of software engineering experience.
  • The ability to identify a problem, rapidly develop a scalable solution and ship it.
  • Ability to lean in and assist team members working on critical or complex technical initiatives.
  • Ability to set the technical direction for a specific project and execute.
  • Expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP etc.)
  • Strength in at least one programming language - Go, Python, Java, Rust.
  • Strong analytical and problem-solving skills.
  • Excellent communication and collaboration skills.
  • Ability to work independently and within a team


Nice to Have:
  • Experience with Temporal and Kubernetes.
  • Experience working directly with hardware vendors.
  • Background in large-scale GPU fleet operations or hyperscale data center environments.


Benefits:
  • Industry competitive pay
  • Restricted Stock Units in a fast-growing, well-funded technology company
  • Health insurance package options that include HDHP and PPO, vision, and dental for you and your dependents
  • Employer contributions to HSA accounts
  • Paid Parental Leave
  • Paid life insurance, short-term and long-term disability
  • Teladoc
  • 401(k) with a 100% match up to 4% of salary
  • Generous paid time off and holiday schedule
  • Cell phone reimbursement
  • Tuition reimbursement
  • Subscription to the Calm app
  • MetLife Legal
  • Company paid commuter benefit; $50 per pay period


Compensation Range

Compensation will be paid in the range of up to $170,000 - $205,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Similar Jobs

More Jobs at Crusoe

More Information Technology Jobs

Find similar Senior Software Engineer (DCIE) jobs: