Member of Technical Staff, Fleet

Mount Thor

$130K — $160K *
US-AnywhereRemote in San Francisco, CA
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Experience managing large production systems essential for inter-team reliance.
  • Proficient in programming languages such as Go, Python, or Rust.
  • Background in distributed systems and control plane technologies is crucial.
  • Familiarity with Linux, macOS, or Unix systems, including boot processes and system performance.
  • Hands-on experience with bare-metal compute management and Kubernetes is preferred.
  • Understanding of integrating node-level software with distributed control systems is necessary.
  • Expertise in utilizing coding agents for software development and operational management.

Responsibilities

  • Develop the strategic roadmap for the fleet control infrastructure.
  • Create distributed control systems and node-level software to manage the fleet's lifecycle.
  • Automate the ingestion of capacity across various Apple hardware models.
  • Link fleet health metrics to workload scheduling for improved performance.
  • Embed agentic development processes into the team's core operations.
  • Instill robust operational practices to monitor health signals and manage incidents.

Benefits

  • Opportunity to work with cutting-edge compute technologies.
  • Emphasis on innovation through agentic engineering techniques.
  • Collaborative team environment with cross-functional responsibilities.
  • A chance to impact the efficiency and reliability of a large-scale compute fleet.
Full Job Description
Mount Thor is hiring a software engineer to build the systems that manage our compute fleet at scale.

The Team

The Fleet team owns the software layer that turns raw compute capacity into a coherent fleet. The team defines how machines enter, operate within, recover, and leave the fleet. It connects physical capacity to workload demand. It establishes the systems and policies that keep the fleet reliable, secure, and efficient as it grows. Fleet determines how quickly new capacity becomes usable and how reliably workloads run. Our goal is to maximize healthy, schedulable capacity and minimize capacity stranded by provisioning failures, hardware faults, or incomplete recovery.

Fleet is also a proving ground for agentic engineering at Mount Thor. We build systems that software agents can inspect and operate safely. Agents should handle routine investigation, change, and recovery. Humans set policy, manage risk, and resolve novel failures.

The team works across hardware, operating systems, networking, scheduling, security, and data center operations.

In this role you will
  • Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle. Lead complex work across teams and systems.
  • Build the distributed control plane and node-level software that manage inventory, configuration, health, and lifecycle state. Make every action safe, observable, auditable, and recoverable.
  • Automate capacity ingestion across Apple hardware generations. This includes provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service.
  • Connect fleet health and capacity to workload scheduling. Improve availability, placement, utilization, recovery time, and the speed at which new capacity reaches production.
  • Make agentic development and operation core to the team. Use coding agents throughout investigation, implementation, testing, and operations. Build interfaces that let software agents inspect state, take safe action, verify results, and escalate exceptions.
  • Establish strong operational practices. Define health signals and service objectives, lead incidents, improve on-call health, and turn failures into lasting system improvements.


You might thrive in this role if you have
  • Built and operated large production systems that other teams depend on.
  • Strong software engineering skills in Go, Python, Rust, or a similar language.
  • Experience with distributed systems, control planes, state machines, controllers, or durable workflows.
  • Strong knowledge of Linux, macOS, or Unix systems. You are comfortable with boot flows, processes, networking, storage, containers, and system performance.
  • Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets.
  • Experience connecting node-level software to distributed control planes or automated operators.
  • Deep experience using coding agents to build production software. You know how to provide the context, tools, tests, and constraints required for reliable results.
  • A track record of leading complex, multi-team infrastructure work from strategy through production.
  • Strong operational judgment. You design for partial failure, safe retries, auditability, and recovery.


Bonus Skills
  • Built fleet-management systems for thousands of machines across multiple sites.
  • Operated Apple Silicon or macOS infrastructure at scale.
  • Built host agents, health daemons, provisioning pipelines, or automated repair systems.
  • Managed scheduling and capacity across several hardware generations.
  • Built infrastructure designed to be operated by software agents, including permissions, validation, rollback, and human escalation.
  • Delivered measurable improvements in availability, utilization, provisioning speed, or recovery time.


Similar Jobs

More Jobs at Mount Thor

More Enterprise Technology Jobs

Find similar Member of Technical Staff, Fleet jobs: