Staff Platform Engineer - Developer Infrastructure

Persona AI

$135K — $160K *
Enterprise Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of experience in production infrastructure for software engineering, specifically in CI/CD or developer platforms.
  • Strong fluency in Linux systems focusing on networking, storage, systemd, and kernel debugging.
  • Hands-on experience with major CI systems and understanding of build and cache layers.
  • Practical experience with C/C++ build systems at scale including cross-compilation and dependency management.
  • Proficient in Infrastructure as Code with tools like Ansible or Terraform.
  • Strong programming skills in Python and Bash for automation purposes.
  • Excellent written communication skills for clear documentation of decisions.

Responsibilities

  • Manage the build graph for a complex C++/Python/Rust/CUDA monorepo.
  • Oversee CI capacity and ensure timely PR feedback through optimized queue management.
  • Guarantee reproducibility of builds from previous commits.
  • Handle versioning and artifact promotion for deployment.
  • Implement safe and efficient rollout processes for robots, including canary testing and rollback strategies.
  • Create reproducible development environments for engineers across devices.
  • Establish observability and monitoring for fleet operations and site infrastructure.

Benefits

  • Flexible work environment with the potential for remote work.
  • Opportunity to work on cutting-edge robotics technology.
  • Collaborative and innovative team atmosphere.
  • Access to professional development resources and training.
  • Comprehensive health and wellness benefits.
Full Job Description
Job Title: Staff Platform Engineer - Developer Infrastructure

Department: Software

Employment Type: Full-Time

Location: Houston, TX

Travel: 10%

About the Role

Our robots are built by a small team that ships fast and deploys onto hardware in customer facilities. Between a laptop and a running humanoid there are: a C++/Python/Rust monorepo, an ARM64 cross-compile matrix, kernel modules and a real-time patch set, container images that have to land on Jetson devices, and a rollout process that cannot brick a robot several timezones away.

Today that path is held together by the engineers who also write control code. Your job is to own it so they don't have to. Success is measured in build times, time-to-first-commit for a new engineer, deployment frequency, and the number of infrastructure problems the rest of the team stops thinking about.

This is not a cloud-only role. Roughly half the surface area is physical: lab networks, robot dev boxes, bench and HIL fixtures, on-site LAN infrastructure, and the mirrors that let a deployment site work with no internet at all.

What You Will Be Doing

Build and CI
  • The build graph for a mixed C++/Python/Rust/CUDA monorepo: involving incremental correctness, remote caching, and cross-compilation for AMD64 dev boxes and ARM64 targets.
  • Our CI capacity: self-hosted runners, GPU and hardware-attached runners, and the queue discipline that keeps PR feedback under ten minutes.
  • Reproducibility: A build from a tagged commit six months from now should produce a bit-identical artifact.

Release and fleet delivery
  • Versioning, artifact promotion, and the container registry / package mirrors that back it.
  • Safe, resumable, bandwidth-aware rollout to robots - staged channels, canary robots, and rollback that works over a bad link.
  • Signing and provenance for anything that lands on a robot.

Developer platform
  • Reproducible dev environments across laptops, shared dev machines, and robots.
  • Self-service tooling so an autonomy engineer can get a branch onto a robot without filing a ticket or learning Kubernetes.
  • Onboarding path: a new engineer builds, tests in sim, and deploys to a bench robot on day one.

Infrastructure and observability
  • Cloud and on-prem compute, storage for multi-TB robot logs, and the training/simulation cluster's operational layer.
  • Fleet observability: metrics, logs, and traces from robot to dashboard.
  • Site infrastructure for deployments: VPN/overlay networking, local mirrors, offline-capable registry authorization.


What We Are Looking For
  • 8+ years operating production infrastructure for a software engineering org, with direct ownership of CI/CD or developer platform work.
  • Deep Linux systems fluency including networking, storage, systemd, kernel and driver debugging.
  • Hands-on ownership of a major CI system (GitHub Actions, GitLab CI, Buildkite, Jenkins) and an understanding of the build and cache layers underneath it.
  • Real experience with a C/C++ build system at scale e.g., Bazel, CMake, or equivalent. Should include cross-compilation and dependency pinning.
  • Infrastructure as code (Ansible, Terraform) and a GitOps mindset: reviewable, reproducible, version-controlled change.
  • Fluent Python and Bash. You automate rather than document a manual procedure.
  • Strong written communication. You leave behind decisions that outlast the conversation.


Bonus Skills
  • Shipping software to embedded or edge Linux targets - ARM64, Jetson, Yocto/custom images, A/B partitions, OTA update systems.
  • Hybrid on-prem plus cloud, and clear judgment about which belongs where.
  • Open-source observability stack (Prometheus, Grafana, OpenTelemetry) and self-hosted services (registries, artifact stores, object storage).
  • Robotics, autonomous vehicles, aerospace, or another hardware-heavy environment where a bad release has physical consequences.
  • Nix, Bazel remote execution, or other tools for reproducible builds at scale.

Similar Jobs

More Jobs at Persona AI

More Enterprise Technology Jobs

Find similar Staff Platform Engineer - Developer Infrastructure jobs: