Fleet Reliability Engineer

DYNA Robotics Inc

$110K — $130K *
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3 to 5 years in robotics, mechatronics, or reliability engineering with experience on fielded hardware.
  • Hands-on skills in electronics repair, including soldering and assembly work.
  • Understanding of actuator mechanics, including wear and failure modes.
  • Knowledge of robot control systems and tuning parameters for deployed behavior.
  • Strong Python skills for data analysis and tooling development.
  • Fluency with actuators, motor control, and calibration processes.
  • Experience with observability tools and version control in git.

Responsibilities

  • Root-cause complex hardware failures across the robot fleet.
  • Physically test and repair suspect components and systems.
  • Design and fabricate mechanical pieces for testing and repair.
  • Translate mechanical data into actionable design feedback.
  • Develop and implement fleet health metrics for monitoring and alerts.
  • Create calibration tools and test tooling for new hardware releases.
  • Manage robot telemetry extract, decoding, and monitoring at scale.
  • Document processes clearly for team handoff and collaboration.

Benefits

  • Flexible working hours to support work-life balance.
  • Opportunities for professional development and skill enhancement.
  • Collaborative work environment with a focus on innovative robotics solutions.
  • Access to advanced technology and tools for hands-on engineering work.
Full Job Description
About the role

We are looking for an engineer to own the health of our humanoid robot fleet, from root-causing a single hardware failure at the bench to building the telemetry and tooling that catches degradation before it becomes downtime.

This is a genuinely hands-on role. The same person who writes the fleet-wide degradation analysis should be comfortable pulling an arm off a robot, probing a suspect driver board, soldering a fix, and designing the bracket that makes the repair stick. You will sit at the intersection of hardware, firmware, controls, and data infrastructure, and you will be the person the team calls when a robot does something nobody can explain.

What You'll Do
  • Root-cause complex failures across the fleet. Actuator dropouts, encoder and calibration drift, thermal events, CAN bus faults, and control-stack bugs, separating genuine hardware degradation from firmware, calibration, and model-side causes.
  • Work hands-on with the hardware. Bench-test suspect actuators and electronics, rework and repair (soldering, harness and connector work), swap and re-commission components, and run physical validation with calipers, test fixtures, and boot and power-cycle rigs you build yourself.
  • Design and build the mechanical pieces. Take test fixtures, jigs, brackets, mounts, and small modifications to fielded robots from concept through fabrication, whether 3D printed or through the shop or a vendor.
  • Turn mechanical evidence into design changes. Reason from your data about backlash, windup, gear wear, preload, and hard-stop behavior, and feed that back as concrete changes. The mechanical team owns major structural design and simulation, so you will work closely with them and with vendor models and joint-limit specs.
  • Build and roll out fleet health metrics. Maintenance indices, hold-current precursors, backlash and deadband tracking, and bus-error monitoring, with alert rules tuned to fire on real faults rather than noise floors.
  • Develop calibration and test tooling. Calibration tools with safety gates, characterization sweeps for new hardware revisions, and pre-merge validation harnesses for firmware and software changes.
  • Wrangle robot telemetry at scale. Extract and decode logs (MCAP, npz, episode archives) from bandwidth-constrained robots, close gaps in the telemetry pipeline, and build the monitoring that should have existed already.
  • Write things down and hand them off. Runbooks, SOPs, incident reports, and handoff packages clean enough that a teammate can pick the work up cold.
  • Review PRs and propose fixes. Across firmware defaults, control code, and deployment configuration, catching interaction bugs between calibration tools, boot checks, and safety gates.


What You'll Bring
  • 3 to 5 years in robotics, mechatronics, hardware test, or reliability engineering, with significant time on fielded hardware rather than lab prototypes.
  • Genuinely hands-on instincts. You are comfortable at the bench with a soldering iron, multimeter, and calipers, and you can tear down, repair, and re-commission electromechanical assemblies without waiting on another team.
  • An understanding of how actuators fail mechanically, including backlash, wear, preload loss, and end-stop damage, and the ability to connect physical evidence to design causes. You will not be doing large structural design or FEA, but you should hold your own in a design review.
  • Working knowledge of robot control: impedance and torque-controlled actuators, force-control parameters, and controller state lifecycles (boot, e-stop, reconnect). Enough to tune deployed behavior and recognize when a "hardware" fault is actually a controls bug. You do not need to design controllers from scratch.
  • Strong Python for data analysis and tooling, plus solid time-series and signal-statistics instincts around noise floors, drift, persistence rules, and spotting when an alert is just a tiny denominator.
  • Fluency with actuators and motor control: quasi-direct-drive systems, encoders, torque and current telemetry, CAN, and calibration and zeroing procedures.
  • Experience with observability stacks (Datadog or similar) and git and PR-based workflows.
  • Rigor about evidence. You distinguish measured from commanded and correlation from cause, and you will not ship a conclusion or set a baseline on contaminated data.
  • A bias toward small, verifiable scope. You deliver the stated goal, flag adjacent issues once, and do not gold-plate.
Bonus points for
  • Experience with teleoperation systems, VR interfaces, or data-collection fleets.
  • Prior work on hardware qualification (EVT, DVT, PVT) or manufacturing and incoming test.
  • Simulation sanity-checking in MuJoCo.
  • A track record of building tools and SOPs that other engineers actually adopt.


Don't let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you're passionate about closing the loop between deployed robots and better models, we want to hear from you, even if you don't check every box.

Similar Jobs

More Jobs at DYNA Robotics Inc

More Technical Services Jobs

Find similar Fleet Reliability Engineer jobs: