Lead Engineer, Issue Management & Triage

Diligent Robotics

$125K — $150K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years in technical or program management roles (engineering or incident management)
  • 3+ years of experience in people management
  • Familiarity with complex systems like robotics and autonomous technologies
  • History of building operational tools and infrastructures
  • Strong programming skills, particularly in Python
  • Experience with data pipelines and telemetry systems
  • Comfortable with tools like Jira, SQL, and Looker

Responsibilities

  • Design and manage systems for issue intake and triage
  • Establish severity frameworks and SLAs for prioritizing issues
  • Develop automation pipelines to improve data processing efficiency
  • Collaborate with engineering to integrate operational tools
  • Translate real-world problems into technical categories for resolution
  • Create classification frameworks to measure robot performance
  • Initiate root cause analysis and promote continuous system improvements

Benefits

  • Remote work flexibility available within the U.S.
  • Opportunity for substantial travel to the Austin HQ during initial onboarding
  • Engagement in cutting-edge robotic technology and problem resolution
  • Chance to work closely with cross-functional teams including engineering and operations
  • Hands-on role with direct contributions to product and system development
Full Job Description
As a Fleet Engineer, you9ll own the reliability and continuous improvement of our deployed robotic fleet - leading hands-on investigations into how and why robots fail in the field, across the mobile base, charging/docking, motion and power, connectivity (modem), and sensor hardware. You9ll combine remote data analysis with bench/lab failure analysis at our Austin HQ, turning field-technician reports and fleet data into clear problem statements, validated root causes, and corrective actions driven to closure with engineering, operations, manufacturing, and vendors.

We are hiring a Lead Engineer, Issue Management & Triage to lead the systems, tooling, and team at the intersection of our Customers, Remote Operations Center (ROC), and Engineering. This is a highly technical, hands-on role focused on building the infrastructure that powers how we detect, triage, diagnose, and resolve issues across a deployed robotic fleet. You will work deeply with Engineering teams to design classification frameworks, build internal tools, and develop automation pipelines that improve reliability at scale.

Location: Austin preferred, Remote possible (U.S.)
Travel: if remote up to ~50% travel to Austin, TX (especially in the your first 90 days)

What You9ll Do:

Own Issue Management & Triage Systems
  • Design and own end-to-end systems for issue intake, triage, and escalation.
  • Define severity frameworks, SLAs, and ensure issues are consistently structured for engineering prioritization.

Build Tools & Automation (Hands-On)
  • Develop automation and pipelines to ingest, process, and classify operational data, reducing manual triage effort.
  • Contribute directly to codebases (Python, backend services) and partner with Engineering on system integrations (logs, telemetry, alerts).

Bridge Operations & Engineering
  • Act as the primary technical interface between the Remote Operations Center (ROC) and Engineering.
  • Translate real-world issues into prioritized, categorized technical problems for resolution alignment.

Performance Measurement & Classification Frameworks
  • Develop systems and taxonomies to systematically measure and classify robot performance, failure modes, and degradation across the fleet.
  • Build dashboards and reporting systems to track trends, severity, and impact.

Root Cause Analysis & Continuous Improvement
  • Establish best practices for Root Cause Analysis (RCA) and identify systemic issues.
  • Drive long-term fixes and create feedback loops to influence improvements in hardware, software, and autonomy.

What We9re Looking For:
  • 7+ years in relevant technical or program management roles (e.g., engineering, incident management)
  • 3+ years of people management
  • Experience with complex, real-world systems (robotics, autonomous/distributed systems, or hardware-software products)
  • Proven track record building operational tools, systems, or infrastructure for workflows

Technical Skills
  • Strong programming experience (Python preferred; backend or data systems experience a plus)
  • Experience with:
    • Data pipelines and telemetry systems
    • Monitoring, alerting, and logging infrastructure
    • Internal tools and automation systems
  • Ability to design scalable systems for classification, prioritization, and workflow automation
  • Familiarity with platforms like Jira, Zendesk, SQL, Looker, Foxglove, or similar

Systems & Product Thinking
  • Strong systems thinker, translating ambiguous operational problems into structured technical solutions
  • Experience defining metrics, taxonomies, and performance frameworks
  • Data-driven approach to prioritization and decision-making

Mindset
  • Hands-on and willing to dive into technical problems when needed
  • Strong ownership and bias toward action
  • Comfortable operating in a fast-paced, scaling environment
  • Passion for improving real-world system performance and reliability

Similar Jobs

More Jobs at Diligent Robotics

More Enterprise Technology Jobs

Find similar Lead Engineer, Issue Management & Triage jobs: