Member of Technical Staff, Fleet

Mount Thor

$150K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in software engineering, ideally with a focus on systems automation.
  • Proficient in programming languages such as Go, Python, or Rust.
  • Hands-on experience with distributed systems and control plane development.
  • Strong understanding of Linux, macOS, or Unix operating systems and their internals.
  • Experience with bare-metal compute environments and large server fleets management.
  • Background in connecting node-level software to distributed systems and automated operations.
  • Proven ability to lead complex infrastructure projects and multi-team collaborations.

Responsibilities

  • Own and drive the technical strategy for fleet operation and lifecycle management.
  • Develop distributed control plane software for managing system inventory and health.
  • Automate processes for ingesting compute capacity across different Apple hardware generations.
  • Connect fleet health with scheduling to enhance operational reliability and performance.
  • Integrate agentic software development practices into the team’s workflow.
  • Define operational targets and improve system reliability through incident management.

Benefits

  • Innovative work environment at the forefront of infrastructure development.
  • Opportunity to work with cutting-edge technologies and complex distributed systems.
  • Collaborative culture that fosters teamwork and multi-disciplinary engagement.
  • Access to learning resources and professional development opportunities.
Full Job Description
Mount Thor is hiring a software engineer to build the systems that manage our compute fleet at scale.

The Team

The Fleet team owns the software layer that turns raw compute capacity into a coherent fleet. The team defines how machines enter, operate within, recover, and leave the fleet. It connects physical capacity to workload demand. It establishes the systems and policies that keep the fleet reliable, secure, and efficient as it grows. Fleet determines how quickly new capacity becomes usable and how reliably workloads run. Our goal is to maximize healthy, schedulable capacity and minimize capacity stranded by provisioning failures, hardware faults, or incomplete recovery.

Fleet is also a proving ground for agentic engineering at Mount Thor. We build systems that software agents can inspect and operate safely. Agents should handle routine investigation, change, and recovery. Humans set policy, manage risk, and resolve novel failures.

The team works across hardware, operating systems, networking, scheduling, security, and data center operations.

In this role you will
  • Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle. Lead complex work across teams and systems.
  • Build the distributed control plane and node-level software that manage inventory, configuration, health, and lifecycle state. Make every action safe, observable, auditable, and recoverable.
  • Automate capacity ingestion across Apple hardware generations. This includes provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service.
  • Connect fleet health and capacity to workload scheduling. Improve availability, placement, utilization, recovery time, and the speed at which new capacity reaches production.
  • Make agentic development and operation core to the team. Use coding agents throughout investigation, implementation, testing, and operations. Build interfaces that let software agents inspect state, take safe action, verify results, and escalate exceptions.
  • Establish strong operational practices. Define health signals and service objectives, lead incidents, improve on-call health, and turn failures into lasting system improvements.


You might thrive in this role if you have
  • Built and operated large production systems that other teams depend on.
  • Strong software engineering skills in Go, Python, Rust, or a similar language.
  • Experience with distributed systems, control planes, state machines, controllers, or durable workflows.
  • Strong knowledge of Linux, macOS, or Unix systems. You are comfortable with boot flows, processes, networking, storage, containers, and system performance.
  • Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets.
  • Experience connecting node-level software to distributed control planes or automated operators.
  • Deep experience using coding agents to build production software. You know how to provide the context, tools, tests, and constraints required for reliable results.
  • A track record of leading complex, multi-team infrastructure work from strategy through production.
  • Strong operational judgment. You design for partial failure, safe retries, auditability, and recovery.


Bonus Skills
  • Built fleet-management systems for thousands of machines across multiple sites.
  • Operated Apple Silicon or macOS infrastructure at scale.
  • Built host agents, health daemons, provisioning pipelines, or automated repair systems.
  • Managed scheduling and capacity across several hardware generations.
  • Built infrastructure designed to be operated by software agents, including permissions, validation, rollback, and human escalation.
  • Delivered measurable improvements in availability, utilization, provisioning speed, or recovery time.


Similar Jobs

More Jobs at Mount Thor

More Information Technology Jobs

Find similar Member of Technical Staff, Fleet jobs: