Member of Technical Staff, Fleet Operations

Mount Thor

$125K — $150K *
US-AnywhereRemote in San Francisco, CA
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in production data centers or large-scale compute infrastructure
  • Strong knowledge of data center systems (racks, power, cooling, networking)
  • Proficiency in macOS, Linux, or Unix with troubleshooting capabilities
  • Experience designing commissioning and maintenance systems for physical infrastructure
  • Strong programming skills in languages like Python, Go, or Bash
  • Experience in building telemetry pipelines and diagnostic tools
  • Ability to lead root-cause analysis and implement reliability improvements

Responsibilities

  • Own the technical direction for data center architecture and operations
  • Design and automate commissioning workflows for new capacity
  • Build observability tools for hardware and environmental telemetry
  • Lead investigations of failures across hardware and software integrations
  • Implement improvements based on failure analyses and insights
  • Design maintenance and hardware lifecycle management systems
  • Lead production incidents and coordinate with cross-functional teams

Benefits

  • On-site engineering role with exposure to high-performance compute systems
  • Opportunity to innovate and improve critical infrastructure processes
  • Access to advanced technology and production environments
  • Engagement with a collaborative, cross-disciplinary engineering team
  • Involvement in continuous improvements using agent-based development practices
Full Job Description
Mount Thor is hiring an engineer to design and improve the data center systems behind our compute fleet. This is an on-site engineering role with direct responsibility for production infrastructure.

The Team

Fleet Operations Engineering designs and operates the physical infrastructure behind Mount Thor's compute fleet. The team owns data center architecture, hardware integration, commissioning, observability, diagnostics, maintenance, and reliability. It works across racks, power, cooling, networks, operating systems, and Apple hardware to bring capacity online and keep it performing in production. The team builds the instrumentation, tools, standards, and automated workflows needed to find failures, shorten recovery time, and improve each new deployment.

Fleet Operations Engineering partners with Fleet at the machine readiness and health boundary: this team establishes and validates those conditions in the data center, while Fleet represents and acts on them through the software control plane.

In this role you will
  • Own the technical direction for data center architecture and operations engineering. Define standards for rack design, power, cooling, cabling, network integration, serviceability, safety, and security.
  • Design and automate the commissioning of new capacity. Build repeatable systems for installation checks, inventory validation, connectivity testing, hardware qualification, and production readiness.
  • Build observability for the physical fleet. Collect and validate hardware, power, thermal, network, and environmental telemetry. Make that data useful for diagnosis, capacity planning, and automated health decisions.
  • Lead complex failure investigations across hardware and software boundaries. Work directly with production systems to isolate issues involving hardware, firmware, macOS, networking, storage, power, or cooling.
  • Turn failures into engineering improvements. Build better diagnostics, test tools, hardware designs, maintenance procedures, and automated workflows.
  • Design the systems behind maintenance and hardware lifecycle management. Improve preventive maintenance, repair, spare planning, vendor escalation, hardware refresh, and secure decommissioning.
  • Make agentic development part of daily engineering. Use agents to analyze telemetry, develop tools, investigate failures, and improve documentation. Build structured data and workflows that agents can use safely.
  • Lead production incidents and planned infrastructure changes. Participate in on-call coverage and coordinate with Fleet, networking, security, data center partners, and hardware vendors.


You might thrive in this role if you have
  • Designed or operated production data center, HPC, cloud, or large-scale compute infrastructure.
  • Strong knowledge of data center systems, including racks, power, cooling, structured cabling, networks, and environmental monitoring.
  • Deep systems knowledge across macOS, Linux, or Unix. You can troubleshoot boot flows, firmware, storage, networking, system performance, and hardware-software interactions.
  • Experience designing commissioning, qualification, diagnostic, or maintenance systems for physical infrastructure.
  • Strong programming skills in Python, Go, Rust, Bash, or a similar language.
  • Experience building telemetry pipelines, dashboards, alerts, or diagnostic tools using metrics, logs, and time-series data.
  • A record of leading root-cause analysis and turning individual failures into broader reliability improvements.
  • Strong operational judgment. You plan changes carefully, validate outcomes, document decisions, and design for safe recovery.
  • Experience using coding agents or other AI tools for engineering, investigation, and data analysis.
  • Comfort working in active data center environments, joining an on-call rotation, and traveling to other sites when needed.


Bonus Skills
  • Operated Apple Silicon or macOS infrastructure at data center scale.
  • Worked with macOS recovery, firmware, secure boot, hardware diagnostics, or automated device restoration.
  • Designed custom racks or data center environments for dense, nontraditional compute hardware.
  • Integrated building, power, environmental, or hardware telemetry through industrial protocols and vendor APIs.
  • Applied hardware reliability methods such as failure analysis, predictive maintenance, component-life tracking, or data-driven spare planning.
  • Built commissioning or diagnostic systems used across multiple data center sites.
  • Designed AI-assisted operational systems with clear permissions, audit trails, validation, and human escalation.


Similar Jobs

More Jobs at Mount Thor

More Information Technology Jobs

Find similar Member of Technical Staff, Fleet Operations jobs: