Decision Engineer, Compute Operations

Fluidstack

• $110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of software engineering experience with production-grade code in Go, Python, and TypeScript
  • Strong full-stack development skills in web applications using frameworks like React or Next.js
  • Hands-on experience with LLM APIs and agentic frameworks
  • Daily use of AI coding agents with demonstrated tools and skills
  • Proven capability to work autonomously, identifying problems and shipping solutions
  • Strong product and systems judgment, comfortable with ambiguity
  • Excellent communication skills for translating technical concepts to non-technical stakeholders

Responsibilities

  • Build decision systems and internal tools to enhance productivity in compute production flow
  • Own tools from conception to production, working closely with users to iterate
  • Design and maintain shared AI infrastructure and build shared skills for the organization
  • Create custom integrations between LLMs and internal systems like Slack and GitHub
  • Stay updated on AI advancements to proactively identify applications

Benefits

  • Competitive total compensation package including cash and equity
  • Health, dental, and vision insurance
  • Retirement plan
  • Generous PTO policy
Full Job Description
The Decision Team

Examples of key problems the team is working on
  • Automate the delivery of gigawatts. Every process that takes AI infrastructure from land to live compute becomes software: schedules, decisions, and todos generated from a live knowledge graph instead of chased by hand.
  • Forward-deploy beside the experts. Product teams sit with quality managers, sourcing leads, and deployment engineers on factory floors and sites, and turn their judgment into systems that reach every unit.
  • Deliver every supercomputer faster than the last. Dozens of concurrent projects feed one graph, so every lesson learned at one site becomes a preventive check at all of them.
Role Scope
  • Build the fleet health system: real-time telemetry and tiered healthchecks on every machine across Kubernetes and bare metal, rolled into one API the whole company trusts to answer "is this machine healthy," with alarms correlated into incidents that reach on-call with a drafted probable cause.
  • Turn repair and RMA into generated work: one tracked flow from failure detection through triage, parts, vendor return, and return to service, where failure thresholds route machines to repair automatically, each production engineer's shift todo list is generated for them, and time to return to service is a number the system reports.
  • Ship hardware qualification as software: burn-in, performance baselining, and new hardware validation composed into rack-level workflows, so bringing thousands of accelerators online is a repeatable run and every machine enters production with its acceptance evidence attached in the graph.
  • Run the facility on the same system as the fleet: the maintenance system for lockout tagout and work orders is live at one site and rolls out to two more, every asset register loads before the first external audit this fall, and the legacy datacenter inventory retires before the next building energizes. You own the asset model, the migration, and the day the old tools switch off.
  • Turn every runbook into a checked procedure: SOPs, training records, and technician qualifications become structured data the customer can audit, and site SLOs, deployment cycle time, and labor ramp report themselves on the dashboards a hyperscaler customer asked for. You work forward-deployed beside production engineers and facility operators, on site and on the rotation, and build what they use the next shift.
What We're Looking For

The below is a starting point. We always make space for exceptional people, so if you don't fit this role exactly,tell us where you would.
  • You've shipped production code in Go, Python, or TypeScript, and you pick up whatever language the problem demands.
  • You've built real features on LLM APIs (OpenAI, Anthropic, or open-weight models), MCP servers, and agentic frameworks.
  • You work daily with AI coding tools like Claude Code and Cursor, and you get agents doing useful work autonomously alongside you.
  • You identify problems, design the solution, and ship it without waiting for direction or approval.
  • You've moved fast under deadline while leaving foundations that other engineers extended after you moved on.
  • You've sat the on-call rotation or worked beside the people who do, and you've turned operational pain into systems that made the pager quieter.
  • Your product taste shows in what you've shipped: interfaces the engineers on the rotation call obvious, and workflows that match how the work actually happens.
  • Bonus: Production engineering or SRE on large GPU fleets. Hardware qualification or burn-in frameworks. BMC, Redfish, or IPMI tooling. CMMS, DCIM, or asset management systems. BMS/EPMS or SCADA. Prometheus and Grafana.

    Benefits:
  • Competitive total compensation package (cash + equity)
  • Health, dental, and vision insurance
  • Retirement plan
  • Generous PTO policy


We are committed to pay equity and transparency.

Similar Jobs

More Jobs at Fluidstack

  • Fluidstack Labs Lead
    $125K — $150K *
    Austin, TX 78745 (Travis County)
    Telecommunications & Hardware
    In-Person
  • People Programs Lead
    $110K — $130K *
    Austin, TX 78745 (Travis County)
    Business Services
    In-Person
  • People Programs Lead
    $120K — $145K *
    San Francisco, CA 94112 (San Francisco County)
    Business Services
    In-Person
  • Accounting Manager
    $110K — $130K *
    San Francisco, CA 94112 (San Francisco County)
    Legal & Accounting
    In-Person
  • People Programs Lead
    $110K — $130K *
    New York, NY 10025 (New York County)
    Business Services
    In-Person

More Information Technology Jobs

Find similar Decision Engineer, Compute Operations jobs: