Fleet Reliability Engineer

Specter

• $120K — $145K *
Transportation
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Strong data and software skills in Python (or Go) and SQL.
  • Hands-on experience building and tuning observability stacks like OpenTelemetry, Grafana, or Prometheus.
  • Experience managing physical or embedded device fleets at scale and understanding hardware failure modes.
  • Proficient in analyzing field telemetry for trends, failure modes, and forecasts.
  • Fluent in databases and data modeling (PostgreSQL or similar); familiarity with infrastructure-as-code tools like Terraform is a plus.
  • Strong inclination towards automation and building mechanisms to reduce manual tasks.
  • Familiarity with reliability/SRE principles and experience with hardware-software integration is a plus.

Responsibilities

  • Own the reliability data pipeline from telemetry aggregation to instrumentation.
  • Reduce observability costs by optimizing tooling expenditures.
  • Instrument fleet data to establish health metrics for reliability evaluation.
  • Verify fleet-wide fixes and isolate genuine failures from transient ones.
  • Automate detection and recovery strategies for repeat failure patterns.
  • Analyze fleet telemetry to identify which cohorts and hardware revisions are trending toward failure.
  • Score reliability initiatives economically to prioritize addressing the most costly issues.

Benefits

  • Comprehensive health and wellness programs.
  • Flexible work hours and remote work options.
  • Professional development opportunities.
  • Collaborative and innovative working environment.
Full Job Description
The RoleWe're hiring a Fleet Reliability Engineer to keep our sensor fleet running in the field by building the data, analytics, and recovery mechanisms that prevent failures from becoming incidents. As we scale toward thousands of sensors, fleet health becomes a data-and-systems problem. This is the proactive, highest-leverage side of reliability: own the telemetry and data pipeline, verify that fixes hold fleet-wide, and turn field signal into cost-weighted decisions about what to fix first. Much of today's operational load is addressable through better instrumentation, alert hygiene, and recovery verification, at little to no field cost.
Responsibilities:

Reliability Data Platform - Primary
  • Own the fleet's reliability data pipeline end to end: telemetry aggregation, storage, and instrumentation.
  • Drive down observability cost - own the tooling spend and cut what we pay for but don't use.
  • Instrument the fleet and own the health metrics that measure reliability.

Proof-of-Recovery & Alert Hygiene
  • Verify that fixes hold fleet-wide, not just on the device that paged.
  • Cut alert noise at the source - separate real failures from self-resolving ones.
  • Turn repeat failure patterns into automated detection and recovery.

Fleet Health & Failure-Mode Analytics
  • Turn fleet telemetry into a live picture of which cohorts, hardware revisions, and firmware versions are trending toward failure, and why.
  • Build the failure-mode analysis that tells engineering what to fix at the source.
  • Own fleet-wide trend and forecasting work, including power and solar planning.

Reliability Economics & Prioritization
  • Score reliability work in dollars - field-trip cost, hardware-return cost, observability spend - and prioritize the most expensive problems first.
  • Set and track the fleet's reliability targets: uptime, offline rate, truck-rolls per sensor-year.
  • Give the team the data to make reliability-versus-cost tradeoffs.
Qualifications:
  • Strong data and software skills - Python (or Go) and SQL - and the ability to own a data pipeline end to end.
  • Hands-on building and tuning observability stacks (OpenTelemetry, Grafana, Prometheus, Datadog, or similar), including their cost.
  • Experience operating physical or embedded device fleets at scale, and reasoning about how hardware fails in the field.
  • Comfortable turning messy field telemetry into trends, failure modes, and forecasts.
  • Fluency with databases and data modeling (PostgreSQL or equivalent); infrastructure-as-code familiarity (Terraform or similar) a plus.
  • Bias toward building mechanisms over doing manual work.
  • Nice to have: reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.
  • Nice to have: experience across the hardware-software boundary - power, connectivity, and physical failure modes.

Similar Jobs

More Jobs at Specter

  • Software Engineer - Embedded Systems
    $120K — $145K *
    San Francisco, CA 94112 (San Francisco County)
    Telecommunications & Hardware
    In-Person
  • Data Operations Engineer
    $110K — $130K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person
  • Architecture
    $130K — $155K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Fleet Reliability Engineer
    $120K — $145K *
    San Francisco, CA 94112 (San Francisco County)
    Transportation
    In-Person
  • Senior Antenna Engineer
    $130K — $155K *
    San Francisco, CA 94112 (San Francisco County)
    Telecommunications & Hardware
    In-Person

More Transportation Jobs

Find similar Fleet Reliability Engineer jobs: