5+ years of software engineering experience, including leadership roles.
Proven experience in building and managing production infrastructure software.
Strong background in fleet management and observability for compute infrastructure.
Solid understanding of distributed systems and Linux fundamentals.
Exceptional coding and debugging skills, capable of hands-on problem-solving.
Experience in setting technical direction and mentoring engineers.
Responsibilities
Lead and develop the Fleet Software team, prioritizing tasks and mentoring engineers.
Set architecture and roadmap for fleet provisioning and configuration across deployments.
Establish safe software and firmware rollout strategies, including validation and handling failures.
Guide design of telemetry and log pipelines for effective data management.
Develop monitoring and analysis tools to enhance fleet performance understanding.
Collaborate with Node Systems team to define hardware management interfaces.
Support customer environments through local monitoring and telemetry analysis.
Translate deployment lessons into engineering priorities and tool decisions.
Benefits
Comprehensive medical, dental, and vision packages with generous premium coverage.
$500 monthly credit for waiving medical benefits.
$2,500 monthly housing subsidy for local employees.
Relocation support for new hires moving to San Jose.
Wellness benefits covering fitness and mental health.
Daily lunch and dinner provided in the office.
Unlimited compute budget with ROI justification.
Full Job Description
Job Summary
We're hiring a Fleet Systems Lead to join our Supercomputing organization. This team builds the software that enables Etched to deploy inference clusters at gigawatt-scale. This role presents an opportunity to shape how frontier inference hardware is deployed, managed, and orchestrated for our customers.
We co-design chips, racks, software, and manufacturing methods so frontier models can run with best-in-class throughput, latency, cost, and power efficiency for both prefill and decode workloads. Fleet Software enables deployments of these systems at scale through configuration management, safe software and firmware rollouts, and software recovery, while partnering with other engineering teams to drive product improvements as we ramp production at unprecedented speed.
The team also builds the monitoring and analysis systems that connect signals across the rack. Engineers and their agents use this data to investigate failures, spot patterns and continually improve the performance and reliability of the inference fleet. We're looking for a leader to drive technical direction while building the best team in the industry. Someone who has built infrastructure and owned it through real deployments in the field. This team's work will shape how Etched scales to a computing platform that customers depend on at massive scale.
Key Responsibilities
Lead and develop the Fleet Software team. Set priorities, hire and mentor engineers, and give team members clear ownership.
Set the architecture and roadmap for fleet provisioning, configuration, updates and recovery across labs and customer deployments.
Establish the team's approach to safe software and firmware rollouts, including validation, staged deployment and handling partial failures.
Guide the design of telemetry and log pipelines, including decisions around storage, retention and query performance.
Lead the development of monitoring, alerting and analysis tools that help engineers and customers understand fleet behavior and investigate failures.
Partner with our Node Systems Systems team to define the interfaces and validated configurations needed to manage hardware across the fleet.
Work with deployment teams to support customer environments, including local monitoring and customer telemetry analysis.
Turn lessons from deployments and incidents into engineering priorities, and decide where to adopt existing tools versus building our own.
Stay involved through design reviews, code contributions and hands-on debugging, particularly on the team's hardest technical problems.
You may be a good fit if you have (Must-have qualifications)
5+ years of software engineering experience, including leading projects or managing engineers.
Experience building infrastructure software and owning it in production, including deployments, incidents and the changes that followed.
A strong background in fleet management, provisioning, configuration or observability for compute infrastructure or connected devices.
Solid distributed systems and Linux fundamentals. You understand how systems behave under load, during upgrades and when components fail.
Strong coding and debugging skills. You can work through a difficult problem with an engineer and contribute to the implementation.
Experience setting technical direction, mentoring engineers and getting work shipped across team boundaries.
Strong candidates may also have experience with (Nice-to-have qualifications)
Telemetry and log pipelines, including storage, retention and query performance.
Bare metal servers, accelerator clusters, robotics or other hardware fleets.
Firmware updates, hardware diagnostics and automated recovery.
Software that runs in customer environments as well as internal infrastructure.
Benefits
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2,500 per month for those living within walking distance of the office
Relocation support for those moving to San Jose (Santana Row)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification