The RoleWe are looking for a Software Engineer to lead the architecture and operation of our fleet management, deployment, and recovery systems. Every device we field operates over a long-range wireless link and must remain fully resilient without physical intervention-making fleet reliability a core distributed systems problem.
We are seeking a strong candidate with hands-on experience building software for physical devices and a clear engineering perspective. The core workflow centers on three main pillars: self-testing devices that continuously report status, reliable software deployment pipelines with verified delivery, and a unified control plane providing real-time operational visibility. In this position, you will be responsible for architecting and implementing these core systems end-to-end.
To support these capabilities, our tech stack spans Rust for systems, C on embedded hardware, NixOS for system images, Python for fleet services, and TypeScript for the control plane. We are looking for an engineer eager to work across all of these layers to ensure our control plane seamlessly reflects and manages device state in the field.
ResponsibilitiesFleet Control Plane- Architect and build the unified control plane providing real-time operational visibility across every device fielded over long-range wireless links.
- Own the underlying data model to reconcile claimed vs. actual device states and ensure seamless state representation across the entire tech stack (TypeScript, Python, NixOS, Rust, C).
- Design targeted execution pathways that enable real-time operational management at the device, site, or customer fleet level.
- Maintain control plane trustworthiness and convergence during partial failures or when significant portions of the fleet are offline or unreachable.
Updates & Rollouts- Architect reliable deployment pipelines that handle over-the-air image and firmware updates with verified delivery over constrained wireless links.
- Implement staged rollout strategies and automatic rollback mechanisms to ensure fleet resilience against flawed updates.
- Establish rigorous post-deploy validation checks on edge hardware before confirming software releases across production fleets.
- Extend the release pipeline that feeds all of it - how releases get cut, versioned, and promoted from dev to production, with regression gates that block on real signal.
- Extend the release pipeline from dev to production with automated regression gates and end-to-end hardware testing harnesses.
Self-Test & Self-Healing- Develop self-testing edge software that continuously measures device health at boot and runtime directly on the hardware.
- Build self-healing fault escalation logic to autonomously resolve device failures in the field without physical intervention.
- Engineer resilient recovery mechanisms-such as service restarts, secondary link failover, and fallback images-that guarantee zero site visits.
- Transform raw device events into actionable telemetry, suppressing noise while automatically handling routine operational issues.
Qualifications- Proven experience writing production code in systems languages (Rust, C/C++, Go, or Python) with a deep focus on fleet reliability, performance, and hardware-level failure modes.
- Hands-on background architecting software for physical devices (embedded, IoT, edge).
- Demonstrated success designing unified fleet control planes and data models to reconcile target versus actual device state across high-scale distributed deployments.
- Strong distributed systems instincts: building for partial failures, managing state convergence over lossy networks, and developing self-healing fault-escalation logic for zero site visits.
- Track record of shipping over-the-air (OTA) update pipelines, staged rollouts, and automatic rollbacks with post-deploy edge verification over constrained wireless links.
- Solid networking fundamentals (DNS, firewalls, VPNs, subnets, secure remote access) and experience operating across the full stack from embedded C/Rust to TypeScript control dashboards.
- Nice to have: Experience with sub-OS environments, boot processes, flash storage, and driver integration.
- Nice to have: Experience with declarative build tools and system image management (NixOS, Yocto, Buildroot) for deterministic edge releases.
- Nice to have: Advanced CI/CD pipeline automation (GitHub Actions, Buildkite) integrated with automated regression gates and physical hardware testing harnesses.
- Pragmatic, bias-toward-action engineering mindset focused on delivering reliable operational control planes today while iterating safely.