About The RoleVolta builds and operates large scale GPU compute infrastructure for AI workloads. A single cluster is tens of thousands of GPUs, hundreds of switches, and tens of thousands of cables. At that size you cannot manage the fabric by hand, and you cannot find out whether a design works by deploying it.
This role builds the model and the tooling that make large topologies tractable: a machine readable source of truth for every site, generated configuration that flows from it, and a simulated fabric where changes are proven before they touch hardware. You will write the software that lets a handful of engineers run fabrics that would otherwise need a room full of people.
What You Will Be Doing- Own Volta's network source of truth: the data model covering sites, racks, devices, interfaces, cabling, addressing, and rail-optimized topology, and keep it authoritative rather than descriptive.
- Generate device configuration from that model for every platform in the estate, so that config is an output of the model and never edited in place.
- Build and operate a simulation environment that reproduces full cluster topologies, and make pre-deployment validation of fabric change a normal step rather than an exception.
- Build the CI pipelines that validate network change: schema and policy checks, generated config diffs, simulated convergence, and reachability and routing assertions before merge.
- Detect and close drift between intended state and device state across sites, and make divergence visible rather than discovered during an incident.
- Automate bring-up verification with the bring-up teams: cable plan generation, LLDP based cabling validation, link quality and error checks, and acceptance test suites that produce a pass or fail against the design.
- Build the tooling that turns a new site from a design document into a provisioned fabric, and shorten how long that takes with each deployment.
- Instrument the fabric: streaming telemetry collection, topology aware metrics, and tooling that lets the team reason about a fabric of this size.
- Write production Python or Go in shared repositories, under the same review, testing, and CI standards as the rest of platform engineering.
- Support incident response and root cause work where modeling, simulation, or config history helps explain what happened.
What You Bring- 4+ years in network automation, infrastructure software, or network engineering with a substantial software component.
- Strong Python in production: testing, packaging, code review, and CI. Our working languages are Python, Go, and Rust.
- Data modeling experience with a network source of truth such as NetBox or Nautobot, including extending the model rather than only consuming it.
- Configuration as code in practice: templated or programmatic generation, declarative and idempotent workflows, and version controlled change.
- Network fundamentals at depth: L2/L3, VLANs, BGP, ECMP, leaf spine design, and overlay protocols. You need to understand what you are modeling.
- Hands-on experience with network simulation or emulation, for example containerlab, vendor virtual appliances, or an equivalent lab automation approach.
- Comfort operating at scale: thousands of endpoints and hundreds of devices, where anything that does not generate or validate automatically does not happen.
- Clear written communication. The model and its tooling are used by people who did not build them.
Nice to Have (But Not Essential)None of these are required. Strength in one or more helps.
- Fluency with AI-assisted development, and interest in scaling agent-assisted workflows across the team (agentic CLI tools, MCP, skills, APIs) to amplify delivery.
- Configuration analysis or formal verification tooling such as Batfish.
- NVIDIA Air, SONiC virtual switch, or Cumulus based lab environments.
- gNMI, OpenConfig, or NETCONF/YANG for configuration and telemetry.
- GPU fabric exposure: RoCE v2, InfiniBand, UFM and its API, or NCCL level performance validation.
- Go or Rust, or interest in moving further toward systems level languages.
- Graph based topology modeling, or experience where topology is queried rather than diagrammed.
- Kubernetes operators or controller patterns, and integrating network state into a control plane.
- Ansible, Nornir, or similar frameworks, with a view on where they stop being the right tool.
- Open source contributions to networking or automation projects.
REQ-77
What We OfferAt Volta, we believe people do their best work when they feel supported, trusted and able to grow. We're building a company where you can make an impact, keep a healthy balance between work and life, and build a career you're proud of.
As a global team, we do our best to provide great benefits wherever you're based. While some benefits vary by country due to local regulations, we believe looking after our people is simply the right thing to do.
- Competitive salary based on the work you do here, not your previous salary
- Equity in Volta, giving you the opportunity to share in the company's long-term success
- Retirement/pension contributions
- Comprehensive health, wellbeing and insurance benefits
- Generous number of vacation days each year
Additional InformationBackground ChecksAll offers of employment at Volta are conditional on the satisfactory completion of pre-employment screening, which includes confirmation of your right to work, verification of your employment history and a criminal record check, where this is permitted by local law. Screening is carried out by Zinc, an accredited third-party provider, after an offer is made and all information is handled confidentially and in accordance with applicable data protection law.