Production Engineer, Facilities

Fluidstack

$100K — $120K *
Manufacturing & Automotive
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in production infrastructure or similar roles
  • Proficiency in Python, Go, or similar languages for automation
  • Deep understanding of observability principles in operational contexts
  • Ability to debug across multiple systems from sensors to dashboards
  • Experience with infrastructure as code and deployment pipelines
  • Familiarity with AI-assisted engineering tools
  • Experience with BMS, EPMS, SCADA, and industrial control systems (bonus for applicants)

Responsibilities

  • Lead production incidents from detection to resolution
  • Own facilities telemetry and alarm platform reliability end-to-end
  • Automate diagnosis and repair workflows for efficiency
  • Manage deployment and operational runtime for services
  • Define production-readiness standards and enforce them
  • Establish operational maturity metrics and engage in service reviews
  • Collaborate effectively with cross-functional teams

Benefits

  • Comprehensive health and wellness coverage
  • Opportunities for professional development and training
  • Flexible work arrangements and work-life balance
  • Collaborative team environment and culture
  • Access to modern tools and technologies for work efficiency
Full Job Description
The Production Engineering Team

Examples of key problems the team is working on
  • Scaling the systems which make the physical plant observable and operable as one fleet. Building the telemetry, alarm, topology, and health systems that allow operators and automations to see the true state of power, cooling, and environmental infrastructure across every site.
  • Turn facility incidents into a closed repair loop. Automate the path from detection and diagnosis through maintenance, remediation, validation, and return to service so failures do not disappear into handoffs between software, engineering, vendors, and site operations.
  • Bring new sites and equipment into production safely at construction speed. Build repeatable readiness gates, commissioning signals, staged deployments, canaries, and rollback mechanisms for a fleet growing by multiple sites at once.
  • Keep operators ahead of power and cooling risk. Build capacity views, safeguards, anomaly detection, service-health reviews, and operational tooling that identify problems before they affect customers.


Role Scope
  • Carry the facilities production on-call pager and lead incidents involving facility software, telemetry, controls integrations, and automation. Diagnose the failure, coordinate the responsible teams, restore service, and drive the systemic fix.
  • Own the production reliability of the facilities telemetry and alarm platform end to end. Build and operate ingestion, storage, APIs, data-quality checks, actionable alerts, retention, backups, failover, and recovery across industrial protocols and site integrations.
  • Turn diagnosis and repair into pipelines rather than procedures. Build Python or Go tooling for fleet-wide debugging, maintenance workflows, automated validation, incident response, and safe return to service.
  • Own production deployment and runtime management for facilities services, including BMS and EPMS integrations, SCADA platforms such as Ignition, virtual PLCs, demand management.
  • Define and enforce production-readiness standards for new sites, equipment, APIs, telemetry integrations, and controls deployments. Build the tests, canaries, release gates, staged promotion, and rollback mechanisms that define what healthy looks like before launch.
  • Own the operational maturity of every in-scope service. Establish SLOs, capacity plans, health dashboards, runbooks, escalation paths, incident drills, and regular service reviews with product owners and partner teams.
  • Partner with Facilities Software Automation, Controls and Design Engineering, Field Engineering, and Facilities Operations. You make the systems these other teams build in and consume reliable, observable, scalable, and supportable in production.


What We're Looking For
  • You have carried a pager for production infrastructure and can run an incident from first alert through restoration, postmortem, and systemic fix.
  • You have written production automation in Python, Go, or a similar language that replaced a manual operational workflow other teams depended on.
  • You understand observability as an operating system, not a collection of dashboards. You have defined meaningful service health, alerts, SLOs, and review rhythms.
  • You debug across system boundaries. You can follow a failure from a physical sensor or controller through an industrial protocol, data pipeline, API, dashboard, and operator workflow.
  • You treat toil as a bug. If a repair or deployment requires repeated manual steps, you build the safe, repeatable path.
  • You are comfortable with infrastructure as code, Kubernetes, GitOps, deployment pipelines, and production data systems.
  • You move toward ambiguous, high-impact failures and build enough domain knowledge to make good decisions quickly.
  • You work effectively with software engineers, controls and design engineers, field teams, vendors, and site operators without blurring ownership.
  • You use modern AI-assisted engineering tools to investigate systems, write and review code, and reduce time from diagnosis to resolution.
  • For senior or lead-level scope, you have set technical direction, built a reliability roadmap, grown engineers, and balanced interrupt-driven operations with sustained engineering delivery.
  • Bonus: Experience with BMS, EPMS, SCADA, Ignition, virtual PLCs, BACnet, Modbus, OPC UA, time-series databases, data center power or cooling, alarm rationalization, repair automation, or industrial control security.


We are committed to pay equity and transparency.

Similar Jobs

More Jobs at Fluidstack

More Manufacturing & Automotive Jobs

Find similar Production Engineer, Facilities jobs: