About This Role:You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane - open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream - along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle. The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. Youll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home - building the fork engine, the guest agent, and the multi-substrate model lifecycle. Were looking for individuals whove built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both.
Responsibilities:- Manage the serving tier for open-weight models: engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline.
- Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle.
- Own the stateful data plane: hosted vector stores for semantic memory and graph stores for knowledge graphs - deployed, backed up, scaled, and recovered, with restores that are tested rather than hoped for.
- Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary.
- Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards
- Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards.
- Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds.
Core Qualifications- 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome - deep Go, C/C++, or Zig with genuine appetite for Rust counts. Async runtimes, memory-safety discipline, and debugging at the syscall boundary should be familiar territory.
- Operated Kubernetes workloads that other people depended on - controllers or operators, scheduling, autoscaling, node lifecycle. Youve been paged, and the experience changed how you build.
- Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution - least privilege, no credentials in the sandbox, audit trails, and human approval on write actions.
- Hands-on experience deploying or operating open-weight LLM serving infrastructure (vLLM/SGLang or similar), including packaging models into reliable, metered production endpoints.
- Performance discipline in distributed systems: you measure before you optimize, and you can tell the story of a latency you killed with the numbers attached.
- Practical depth in some of our core stack - Rust (tokio), Python/FastAPI, Kubernetes operators (controller-runtime/Kubebuilder/CRDs), KEDA, Karpenter, GPU device plugins/DRA - with a genuine willingness to research your way into the rest.
- Familiarity with isolation technology (Firecracker, Kata, gVisor, or comparable), secrets management and egress control (Vault/KMS-class), and hosting stateful systems (vector stores like pgvector/Qdrant, graph stores like Neo4j) with backup and failover discipline.
- Comfort operating across cloud and GPU substrates - AWS/Azure/GCP plus managed GPU clouds - using infrastructure-as-code (Terraform/OpenTofu, Bicep) and observability tooling (OpenTelemetry).
- Bachelors Degree; Masters is a plus
Why AZX! - Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.
- Competitive early-stage startup compensation (based on capabilities, experience, and location)
- Bonus eligibility
- Health insurance with meaningful coverage for dependents
- Flexible paid time off
- Equity
- Fully remote culture with a cluster of teammates in Seattle
Additional Information:- Must be able to travel 2x/year for company summits
- Applicants must be currently authorized to work in the United States on a full-time basis.
- We are unable to sponsor or take over sponsorship of employment visas at this time.
- Please note that our interview process includes a written take-home assignment followed by a live two-hour technical session with our engineering team, so if that format isnt a good fit, wed ask that you not apply
- Please only apply to a maximum of 2 roles at a time, any applicants who apply to more then 2 roles within a 6 month period will automatically be disqualified
Next Steps:If this job sounds like a great fit but you dont check
ALL of these qualification boxes, wed still love to hear from you!