Fleet Response Engineer

Rhoda AI

$110K — $130K *
Technical Services
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in technical roles involving debugging of complex systems
  • Proficiency in Linux command line and system data analysis
  • Expertise in coding with Python and/or C/C++
  • Strong communication skills, especially during high-pressure situations
  • Available for on-call duty, including nights and weekends as needed

Responsibilities

  • Monitor and prioritize robot faults and alerts based on operational impact
  • Serve as the primary engineering point of contact for escalations from field teams
  • Debug hardware and software issues through thorough analysis and testing
  • Drive fast incident resolutions to optimize fleet uptime
  • Coordinate with engineering teams for resolution of complex problems
  • Communicate status updates effectively throughout the incident resolution process
  • Conduct root-cause analyses and track corrective actions to improve system reliability

Benefits

  • Opportunity to work with cutting-edge robotics technology
  • Be part of a dynamic environment requiring creative problem-solving
  • Engage in ongoing training and professional development opportunities
  • Collaborative team culture focusing on innovation and reliability
  • Chance to impact the product lifecycle through direct feedback to engineering teams
Full Job Description
The Role

You'll be the first engineer in the loop when something goes wrong on a deployed robot. When the field or operations team hits an issue they can't resolve, they come to you. You'll triage incoming faults and alerts, debug across the full stack: hardware, embedded, software, sensors, networking, to find the fastest safe path back to operation, and pull in the right domain experts when an issue runs deep. You'll own incidents from first alert through recovery and root cause, and turn what you learn into better monitoring, better runbooks, and a more reliable fleet.
What You'll Do
Triage and incident response
• Monitor incoming faults, alarms, and alerts from robots deployed at customer sites and in internal testing; triage and prioritize by severity and operational impact
• Act as the first engineering responder and point of contact for real-time escalations from field technicians and operations
• Debug complex issues across subsystems through log analysis, telemetry review, and reproduction testing
• Drive incidents to fast resolution or mitigation to restore operation and maximize fleet uptime
Escalation and coordination
• Escalate to and coordinate with domain engineering teams (AI, robot software, cloud, hardware) to drive resolution
• Own the on-call rotation and paging, and keep the escalation process clear and current
• Communicate status to operations, engineering, and customer-facing teams throughout an incident
• Own each incident through to confirmed recovery and a clean handoff
Root cause and reliability
• Perform root-cause analysis - identify contributing factors, themes, and corrective actions - and track follow-ups to closure
• Document investigations and fixes in a shared knowledge base and troubleshooting guide so future issues resolve faster
• Feed field insights back to engineering to improve product reliability and issue detection
• Track reliability trends and flag recurring or systemic failures
Tooling and prevention
• Build (or spec, with the software team) the monitoring, alerting, and diagnostic tooling that catches issues fleet-wide
• Establish diagnostic procedures that let technicians and operators self-serve common issues
• Continuously reduce manual, repetitive response work through automation
Required Qualifications
• Strong systematic debugging of complex systems - the ability to reason from symptoms to root cause and reverse-engineer unexpected behavior
• Comfortable in Linux and the command line, and reading logs, telemetry, and system data
• Coding proficiency in Python and/or C/C++ (Bash a plus)
• Clear written and verbal communication, and calm judgment under pressure
• Willingness to take part in an on-call rotation, including some nights and weekends as the fleet grows
Preferred Qualifications
• 3+ years working with robotic, autonomous, automotive, aerospace, or industrial-automation systems in an engineering, reliability, or support capacity
• Familiarity with ROS/ROS2 and robot subsystems (sensors, actuators, perception, networking)
• Prior on-call, site-reliability, or release-engineering experience
• Experience building tools that help others debug, and a track record of solving unusual bugs
• Comfort with hardware and hardware-software interfaces; willingness to travel occasionally to deployment sites

Similar Jobs

More Jobs at Rhoda AI

More Technical Services Jobs

Find similar Fleet Response Engineer jobs: