Bachelor's or Master's in computer science, electrical engineering, robotics, or related field.
1-3 years of fast-growing experience in platform/infrastructure or robotics systems engineering.
Hands-on experience with diagnosing and resolving multi-layer system issues.
Ability to communicate complex technical concepts to non-technical audiences.
Familiarity with CAN, cameras, sensors, or real-time systems is a plus.
Responsibilities
Ensure continuous operation of robots and fix issues at the root cause.
Maintain stable hardware interfaces to simplify upstream development.
Manage observability and data pipelines between robots and the cloud.
Enable internal teams by resolving hardware and software blockages swiftly.
Facilitate direct communication with customers to achieve successful deployments.
Benefits
Flexible working hours to accommodate cross-timezone collaboration.
Opportunities for rapid professional growth with exposure to multiple engineering layers.
Work closely with hardware and deployment teams in dynamic environments.
Full Job Description
You'll be the reason Anvil's robots keep running when nobody's watching - and the person who can explain why, in plain language, to both leadership and the customer standing next to a frozen arm.
The situation you're walking into:
Anvil ships real robots to real customers, and each one depends on a runtime platform that currently has no single owner - reliability work happens, but it's not anyone's full-time job.
The same engineers who should be heads-down on ML, controls, and manufacturing keep getting pulled into platform fires: a broken bring-up jig, a factory test that won't run, a demo config that needs to work by tomorrow.
Most "it doesn't work" reports actually live at the boundary between hardware, ROS2 middleware, and application code - and the root cause is rarely in the layer where the symptom shows up.
This role is deeply communicative in two directions that have nothing to do with writing code: keeping Anvil's leadership informed on what's happening in robot learning and what it means for product and customers, in plain language; and working directly with deployment customers to understand their constraints and get them to a working outcome. Both audiences are non-technical relative to you, and making them smarter is part of the job, not a distraction from it.
What you'll own:
Reliability and root cause: keeping robots running 24/7, and fixing hardware, ROS2, and infrastructure issues at the actual root cause, not with a patch that just moves the symptom.
Clean, stable hardware interfaces - APIs between sensors/actuators and the application layers above them- so upstream teams don't have to think about the hardware boundary.
Observability and cloud: the logging, metrics, debugging tooling, and data pipelines that move information between robots and the cloud.
Internal enablement: unblocking ML, controls, and manufacturing quickly - factory testing workflows, experimental URDF or hardware branches, photoshoot/demo configurations, hardware bring-up and PCIe validation, and generally "closing the loop" on last-mile hardware+software problems.
What the first 100 days look like:
By day 30: ramped on the runtime platform architecture, the ROS2/hardware stack, and current known issues. Has personally closed at least one internal-team unblock (factory testing, bring-up, or a demo config) and shadowed a customer debugging session end to end.
By day 60: has shipped a first real reliability fix or platform improvement that measurably reduces recurring internal escalations. Is the first line of response for at least one class of hardware/software boundary issue.
By day 100: internal teams (ML, controls, manufacturing) are not waiting on platform support for routine needs. Is trusted by leadership as the plain-language translator on robot learning progress, and by customers as a credible technical point of contact.
Who you are:
You care about systems that actually work in the real world, not just clean abstractions on paper.
You're wired toward root cause analysis, not durable-looking patches that just address the symptom.
You move fast: bugs resolved in days, not weeks; features shipped in weeks, not months.
You're comfortable working across boundaries - hardware, middleware, and application layers - rather than staying in one layer.
You prefer incremental improvements over large rewrites.
You're pragmatic: you choose the solution that works reliably over the one that's theoretically elegant.
You're comfortable debugging when the problem is unclear, the logs are incomplete, and the issue comes from an interaction between multiple complex systems.
You can genuinely explain deep technical reality in plain language to people who don't share your technical depth - Anvil's leadership and Anvil's customers alike.
Bonus: familiarity with CAN, cameras, sensors, or real-time systems; experience with containers, CI/CD, and cloud pipelines; exposure to Physical AI models, robot control systems, or computer vision.
Based in or willing to relocate to Taipei or San Francisco - we're hiring for this role in both locations: one seat next to the hardware team and production line in Taipei (strong China-based candidates are also in scope for this seat), one alongside leadership and the demo room in San Francisco - with flexible hours for cross-timezone collaboration when a deployment or debugging issue is urgent.
Education & experience:
Bachelor's or Master's in computer science, electrical engineering, robotics, or a related field. A PhD isn't required or expected - this is a build-it-and-keep-it-running role, not a research role.
Years matter less to us than trajectory. A typical req for a role like this might ask for 3-5 years; we're looking for someone with 1-3 unusually fast-growing years (or 1-2 on a steep curve) spent working directly alongside a senior platform/infrastructure or robotics systems engineer - someone who watched how real cross-layer root-causing happens, closed a growing share of real issues personally, and is ready to carry meaningfully more of that responsibility themselves.
What this role is not:
Not a feature-building software role scoped to a single service. The job is keeping the entire system working continuously, predictably, and debuggably - across layers, not within one.
Not a research role in ML, perception, or controls. You enable those teams and translate their work; you don't own their algorithms.
Not a role for someone who wants to stay in one layer - just ROS2, just hardware, or just cloud. The real work lives at the boundaries between them.
Not a scripted, decision-tree support role. Every internal or customer issue needs real root-cause judgment, not a runbook lookup.