The Data Center Operations TeamThe Data Center Operations team keeps Fluidstack's AI data centers running around the clock, operating the mechanical, electrical, and controls systems that keep gigawatt-scale infrastructure online for frontier AI workloads.Examples of key problems the team is working on:
- Operate live chiller plants, CRAC/CRAH units, and direct-to-chip liquid cooling loops at densities where minutes of cooling loss mean thermal shutdown of GPU clusters.
- Build preventive and predictive maintenance programs that scale across new sites without adding downtime or headcount linearly.
- Take mechanical systems from commissioning handover to steady-state operations on compressed schedules, closing punch-list and performance gaps before they become incidents.
- Run root-cause analysis on mechanical failures and feed corrective actions back into design, operating procedures, and spare parts strategy.
Role Scope- Own mechanical operations on shift as the top escalation point for chiller plant, CRAC/CRAH, air handler, pump, and cooling loop failures in a live, always-on facility.
- Run preventive maintenance execution end to end: scheduling, vendor supervision, lockout/tagout, and QA of completed work, with records rigorous enough to survive a customer or owner audit.
- Lead root-cause analysis for mechanical incidents and drive corrective actions to closure, including procedure updates, setpoint changes, and spare parts adjustments.
- Build and maintain operating procedures (SOPs, MOPs, EOPs) for mechanical systems, and drill them so any operator on shift can execute under failure conditions.
- Drive BMS alarm hygiene and trend analysis so degrading equipment gets caught and fixed before it trips, not after.
- Close out commissioning-to-operations handovers for new mechanical plant, verifying performance against design before accepting systems into the maintenance program.
What We're Looking ForThe below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly, tell us where you would.- 8+ years of hands-on mechanical or critical facilities operations in data centers, with time at a hyperscale operator, colocation provider, or cloud data center operations team.
- Deep operating knowledge of chillers, cooling towers, CRAC/CRAH units, air handlers, pumps, and hydronic loops, including startup, failover, and recovery under load.
- Fluency in BMS platforms: reading trends, tuning alarms, and diagnosing faults from controls data rather than waiting for equipment to fail.
- A track record owning root-cause analysis and preventive maintenance programs at scale, not just executing tickets someone else wrote.
- Experience as the senior mechanical escalation point on shift, making calls during outages when the runbook runs out.
- Willingness to work rotating shifts and carry on-call responsibility for mechanical outages.
- You operate calmly in high-uptime, fast-scaling environments where the plant you ran last quarter is half the size of the one you run now.
- Bonus: experience with direct-to-chip liquid cooling and CDUs, CFC/HCFC refrigerant certification (EPA 608 Universal or equivalent), or a role in commissioning handover for new-build mechanical plant.
We are committed to pay equity and transparency.