AWS IoT is building the infrastructure to deliver AI from the cloud to the physical world. Our Physical AI team is developing services that provision, deploy, monitor, and secure AI software across every processor inside autonomous machines including robots, vehicles, drones, agricultural equipment, and industrial workcells operating in environments with limited or no cloud connectivity.
We are looking for a Software Development Engineer II to own and drive the design and delivery of core subsystems in our edge-first fleet management platform for Physical AI. You will design and build systems that operate reliably without cloud connectivity handling offline provisioning, state reconciliation on reconnect, staged fleet-wide rollouts, atomic rollback across multi-processor machines, and secure over-the-air (OTA) delivery of AI models and software to heterogeneous processor topologies. Your systems must be correct when disconnected, consistent when reconnected, and resilient when partially connected. You will also develop cloud-side services that orchestrate, monitor, and coordinate with fleets of edge devices, ensuring seamless interaction between cloud control planes and disconnected or intermittently connected hardware. You will write high-performance code, working across the hardware-software boundary to debug issues that span operating systems, runtimes, networks, and physical devices. You will independently own substantial subsystems end to end from writing the design document through implementation, testing, and production operation.
You will also mentor junior engineers, raise the bar on design quality, and help shape the technical direction of the platform. This is an opportunity to build foundational infrastructure for a new category of AWS services delivering intelligence to the physical world at scale.
Key job responsibilities
• Design, build, and operate core subsystems of an edge-first fleet management platform, spanning offline provisioning, state reconciliation on reconnect, staged fleet-wide rollouts, and atomic rollback across multi-processor machines
• Develop cloud-side services that orchestrate, monitor, and coordinate with fleets of edge devices, ensuring reliable interaction between cloud control planes and disconnected or intermittently connected hardware
• Build secure over-the-air (OTA) delivery pipelines for AI models and software across heterogeneous processor topologies
• Write and review high-performance, production-quality code, and debug issues that span operating systems, runtimes, networks, and physical devices
• Author design documents for new subsystems and drive them independently from proposal through implementation, testing, and production operation
• Define and uphold correctness guarantees for disconnected, reconnecting, and partially connected system states
• Participate in on-call rotation, troubleshoot production issues, and drive root-cause fixes for the platform
• Mentor junior engineers and raise the engineering bar through code review, design review, and technical guidance
• Collaborate with product, hardware, and adjacent platform teams to define requirements and integration points for new fleet management capabilities
• Contribute to the technical roadmap and architectural direction of the platform as it scales to new device types and processor topologies
A day in the life
Your day starts by checking dashboards - on-call weeks bring pages like an OTA rollback spike; sprint weeks lead into standup with a feature update. Mid-morning often means design review, debating edge cases like reconciling conflicting state after a multi-day disconnect. Sprint weeks give focused coding blocks - building rollback logic across processors, pairing on race conditions. On-call weeks fill that time with live debugging across hardware and software, plus maintenance like patching dependencies. Afternoons bring CR reviews and mentoring a junior engineer through root-causing a bug. You close by updating Taskei and reviewing a peer's design doc.
BASIC QUALIFICATIONS
- 3+ years of non-internship professional software development experience
- Bachelor's degree or equivalent
- Knowledge of professional software engineering & best practices for full software development life cycle, including coding standards, software architectures, code reviews, source control management, continuous deployments, testing, and operational excellence
- • Strong Linux and systems-level skills including production debugging across the hardware-software boundary
- • Practical distributed-systems experience including state machines, idempotency, reconciliation, retries, and persistence
- • Ability to design offline-first and reconnect-and-reconcile behavior for edge systems operating without guaranteed connectivity
- • Experience developing cloud-side services that orchestrate, monitor, and coordinate with fleets of edge devices
- • Track record of independently owning a substantial subsystem end to end from design through production
- • Ability to author solid design documents and make sound architectural decisions without being handed the architecture
- • Experience with fleet management or OTA update systems including staged rollout, rollback, pause-resume, and dependency handling
PREFERRED QUALIFICATIONS
- • Experience with containers at the edge (containerd, K3s, KubeEdge) and judgment on where full Kubernetes is inappropriate for constrained environments
- • Hands-on experience with heterogeneous compute platforms (x86, ARM, GPU, NPU) and accelerator or model versioning
- • Exposure to robotics frameworks such as ROS 2, DDS middleware, NVIDIA Jetson, or Isaac
- • Experience with secure OTA frameworks (TUF, Uptane) and hardware security primitives (TPM, HSM, secure boot)
- • Background in reliability-critical domains such as automotive, autonomous vehicles, industrial IoT, avionics, or telecom edge
- • Experience with large artifact or model distribution including caching strategies and bandwidth optimization
- • Familiarity with AWS IoT, AWS IoT Greengrass, or cloud-based device management services
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Seattle - 143,700.00 - 194,400.00 USD annually