Founding Engineer (Physical AI Infrastructure)

a16z speedrun

$160K — $200K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in building or operating large-scale distributed systems with a focus on reliability.
  • Strong knowledge of Kubernetes, infrastructure as code, networking, storage, security, and cloud architecture.
  • Experience with GPU infrastructure, distributed training, and high-performance computing (HPC) principles.
  • Proficiency in handling high-volume data including streaming, time series, and multimodal datasets.
  • Skilled in programming systems software with languages like Rust, Go, or Python.
  • Ability to troubleshoot across various layers of the tech stack including application, container, and network.
  • A balanced perspective on internal platform architecture and user experience.

Responsibilities

  • Build a unified execution platform for simulation, evaluation, and model training.
  • Design and orchestrate workload scheduling across diverse environments including GPUs and Kubernetes clusters.
  • Develop reliable systems for managing long-running robotics workloads, including checkpointing and scaling.
  • Create robust data infrastructure for various sensor data and telemetry streams.
  • Design storage and indexing systems for large multimodal datasets and ensure data movement efficiency.
  • Develop APIs, SDKs, and workflows for seamless access to infrastructure for engineers and researchers.
  • Oversee platform reliability, incident response, security, and cost efficiency across all environments.

Benefits

  • Opportunity to work with cutting-edge physical AI technology and influential industry trends.
  • Collaborate with a skilled team focused on innovative solutions in AI and robotics.
  • Flexible working environment supportive of both cloud and on-premises infrastructures.
  • Exposure to a dynamic range of technologies, from edge computing to high-performance computing.
  • Possibility to influence and shape the developer experience for a diverse user base.
Full Job Description
The role:

Build the control plane and data plane for physical AI.

Munari must run training, simulation, evaluation, data processing, and inference workloads across public cloud, customer-owned infrastructure, robotics labs, and fleets of edge devices. It must handle GPUs, enormous multimodal datasets, unreliable connectivity, long-running workloads, and machines that cannot simply be restarted whenever something goes wrong.

This is not a conventional DevOps role and it is not a YAML-only role. You will write the systems software that makes physical AI infrastructure feel closer to a programmable platform than a pile of bespoke operations.

What you will work on:
  • Build a unified execution platform for simulation, policy evaluation, model training, data processing, and deployment.
  • Design workload scheduling and orchestration across CPUs, GPUs, Kubernetes clusters, bare metal, on-prem environments, and edge devices.
  • Build reliable systems for checkpointing, retrying, resuming, scaling, and observing long-running robotics workloads.
  • Create the data infrastructure for video, audio, sensor streams, telemetry, trajectories, policy traces, incidents, and human interventions.
  • Design storage, indexing, lineage, replay, retention, and movement of large multimodal datasets.
  • Build the APIs, SDKs, CLIs, and deployment workflows that make the underlying infrastructure simple for robotics engineers and researchers.
  • Own platform reliability, observability, capacity management, incident response, security, tenancy, and cost efficiency.
  • Support disconnected, bandwidth-constrained, private, and potentially air-gapped customer environments.
  • Connect cloud-side infrastructure to the Munari edge runtime and make the entire system operable as one platform.
You may be a strong fit if:
  • You have built or operated large-scale distributed systems where reliability genuinely mattered.
  • You have strong experience with Kubernetes, infrastructure as code, networking, storage, security, and cloud architecture.
  • You understand GPU infrastructure, distributed training, batch systems, MLOps, or high-performance computing.
  • You have worked with high-volume streaming, time-series, image, video, or scientific data.
  • You write production software in Rust, Go, Python, or a similar systems language rather than treating infrastructure as configuration alone.
  • You are comfortable debugging across an application, container, scheduler, network, host, GPU, and storage system.
  • You care equally about the internal architecture and the developer experience exposed to the customer.

Experience with NVIDIA infrastructure, Slurm, Ray, Kueue, Temporal, Kafka, NATS, ClickHouse, Parquet, object storage, multi-cluster Kubernetes, or edge fleet management is useful but not mandatory.

Similar Jobs

More Jobs at a16z speedrun

  • Member of Technical Staff, Research
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Consumer Technology
    In-Person
  • Member of Technical Staff, Research
    $150K — $180K *
    New York City, NY 10025 (New York County)
    Consumer Technology
    In-Person
  • Member of Technical Staff
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • AI Engineer
    $130K — $160K *
    Seattle, WA 98115 (King County)
    Information Technology
    In-Person
  • AI Engineer
    $150K — $180K *
    San Francisco, CA 94112 (San Francisco County)
    Information Technology
    In-Person

More Enterprise Technology Jobs

Find similar Founding Engineer (Physical AI Infrastructure) jobs: