About the Role:Crusoe is a vertically integrated AI Factory company with a mission to accelerate the abundance of energy and intelligence. Our competitive advantage, "Speed is the only moat," is directly tied to our ability to rapidly design, manufacture, and deploy our own modular power and compute infrastructure.
We are seeking a Sustaining Engineering Leader to own how Spark performs after it is deployed. Product Management defines the next generation. New Product Introduction delivers it into production. You own each generation from the moment it leaves ramp: fleet reliability, field failures, configuration, obsolescence, and the design changes that keep deployed units performing for their full service life.
You do not set the roadmap and you do not run the factory. You own the installed base, and you are the evidence source Product Management and NPI both depend on. You will start as the single-threaded owner of this function and build the team as the fleet grows.
First Year MandateSpark deployments begin in early 2027. You will build the system before there is a fleet to run it on.
- Stand up FRACAS, the root cause evidence standard, and the Failure Review Board.
- Establish the as-built configuration baseline before there are units to reconcile.
- Define spares, service intervals, and repair versus replace policy for the first generation.
- Participate in the production release and ramp readiness reviews for the first deploying generation as the receiving organization.
What You'll Be Working On:I. Fleet Ownership & Field Issues- Transfer of Ownership: Take technical ownership of each generation once it completes production ramp, through a defined handover from New Product Introduction covering the released configuration, open issues, and known risks.
- Fleet Reliability: Own failure rate by subsystem, mean time between failures, and repeat failure rate, and define the telemetry every unit must report to make them measurable. Data Center Facility Operations owns uptime and time to repair.
- Closed Loop Corrective Action: Run FRACAS and chair the Failure Review Board. Close every failure with a root cause, a verified fix, and proof it worked. A unit that fails while built within released tolerances is a design issue; attribution unresolved after ten business days proceeds as a Spark-funded fix and escalates to the Spark GM.
- Containment: For safety or fleet-wide risk, contain first and settle attribution after. Own the retrofit decision and its sequencing.
II. Change, Configuration & Obsolescence- Post-Release Change Control: Own engineering change for generations in the field, with cost and fleet impact quantified before approval. New Product Introduction owns change control up to production release; you own it after.
- Configuration Control: Own the configuration baseline and define what Operations records. Reconcile as-maintained against baseline so a change targets exactly the units that need it.
- Obsolescence: Run a proactive DMSMS program. Model last time buys against fleet demand and qualify alternates before a shortage forces the choice.
- Retrofit & Maintenance: Issue engineering change packages, kit definitions, preventive maintenance intervals, and acceptance criteria. Operations executes and returns the records; you verify effectiveness.
III. Feeding Product Management and NPI- Requirements into Product: Convert fleet evidence into quantified design requirements. Product Management decides what enters the PRD.
- Qualification Requirements into NPI: Supply field-derived requirements and acceptance criteria, so a failure the fleet has already seen must be designed out and proven before the next generation is released.
- Total Cost of Ownership: Own the reliability and service inputs to Spark unit economics.
- Serviceability Advocacy: Represent serviceability and maintainability in design reviews. You hold no approval authority; you make the case with fleet data.
What You'll Bring to the Team:Education- Required: Bachelor's degree in Mechanical, Electrical, Reliability, or a closely related engineering discipline.
Experience- 10 to 15 years in sustaining, reliability, or product support engineering for deployed capital equipment or infrastructure hardware.
- Direct ownership of a fielded installed base, with field data converted into design change that measurably reduced failure rate or service cost.
- Obsolescence and lifecycle management on a product whose service life exceeds that of the components inside it.
- Experience standing a function up from nothing, including the processes and the partner agreements behind it.
Technical Depth- Reliability: FRACAS, structured root cause analysis, and mean time between failures and population failure rate analysis.
- Change & Configuration: PLM, change order workflow, as-built control, and serial level effectivity.
- Obsolescence: DMSMS practice, end of life monitoring, last time buy modeling, and alternate part qualification.
- Systems Fluency: Electrical, mechanical, and thermal command sufficient to adjudicate root cause on a modular power and compute product, including liquid cooling.
- Service Economics: Spares, service cost per unit, and the link between availability and revenue.
Additional Qualifications- Experience with modular, containerized, or prefabricated infrastructure products in the field.
- Familiarity with liquid cooled data center hardware, including CDU and rack level cooling interfaces.
- Mountaineer Spirit: The persistence to chase a root cause past the easy answer, and the judgment to know when a fleet-wide fix is worth its disruption.
Benefits:- Competitive compensation and equity packages
- Restricted Stock Units
- Paid time off, paid holidays & leave of absence programs
- Comprehensive health, dental & vision insurance
- Employer contributions to HSA account
- Paid parental leave
- Paid life insurance, short-term and long-term disability
- Professional development & tuition reimbursement
- Mental health & wellness support
- Commuter benefits (parking & transit)
- Cell phone stipend
- 401(k) Retirement plan with company match up to 4% of salary
- Volunteer time off
- Global travel insurance & emergency assistance
- Daily meals allowance
- Additional perks & programs specific to location
Compensation RangeCompensation will be paid in the range of up to $225,000-$255,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicant's knowledge, education, and abilities, as well as internal equity and alignment with market data.