Senior Software Engineer (AI Inference & Runtime Platform)

AZX

$135K — $160K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years experience building production systems in Rust or equivalent systems languages (Go/C/C++/Zig) with a focus on async runtimes and memory safety.
  • Proven experience managing Kubernetes workloads, including autoscaling and node lifecycle management, with hands-on problem-solving.
  • Expertise in threat modeling isolation boundaries for untrusted guest execution, enforcing security best practices.
  • Experience deploying or operating infrastructure for open-weight LLM serving, ensuring reliable model delivery.
  • Strong background in performance measurement for distributed systems, emphasizing data-driven optimization.
  • Proficient in core stack technologies including Rust (tokio), Python/FastAPI, and Kubernetes operators.
  • Familiarity with isolation technologies and stateful system management, including backups and recovery.

Responsibilities

  • Oversee the serving tier for open-weight models, handling deployment, configuration, and performance metrics.
  • Administer Kubernetes for inference and sandbox workloads, focusing on GPU scheduling and lifecycle management.
  • Manage stateful data systems, ensuring proper scaling, backups, and reliable recovery processes.
  • Supervise the sandbox runtime and control processes to ensure security and compliance during execution.
  • Direct the development of FastAPI services and oversee infrastructure management tools like Terraform and OpenTofu.
  • Collaborate with backend services, providing troubleshooting support and maintaining observability tools.
  • Ensure open-source compliance and quality through proper documentation, version control, and build processes.

Benefits

  • Be part of a mission-driven company focused on AI advancements in critical industries.
  • Competitive compensation with potential bonuses and equity participation.
  • Comprehensive health insurance plan including dependent coverage.
  • Generous flexible paid time off policy.
  • Fully remote work culture with occasional team summits.
Full Job Description
About This Role:

You will be responsible for owning the layer where AI work physically happens: the machines, the isolation boundary, and the models running on them. This role anchors on two systems. The first is our inference control plane - open-weight models and custom task-model zoos, hosted and operated across managed GPU clouds and customer-managed Kubernetes clusters, with scale-to-zero economics, cold-start discipline, and per-token cost accounting that stays correct even when a client disconnects mid-stream - along with the Kubernetes layer those workloads live on: operators, autoscaling, node lifecycle. The second is our agent-sandboxing platform: hardware-isolated microVMs for running untrusted, agent-generated code securely and compliantly by construction, where agents operate with least privilege, never see a credential, and a human gates anything that writes to a system of record. Youll write Rust in the morning, a Kubernetes controller after lunch, and a FastAPI control-plane endpoint before you go home - building the fork engine, the guest agent, and the multi-substrate model lifecycle. Were looking for individuals whove built this class of stack (an inference-serving or serverless-GPU platform), operated it hard at scale, or ideally both.

Responsibilities:
  • Manage the serving tier for open-weight models: engine deployment and configuration, cold-start strategy, per-model SLOs, and upgrade/canary discipline.
  • Administer the Kubernetes layer for inference and sandbox workloads: operators and CRDs, autoscaling (KEDA/Karpenter-class), GPU scheduling and sharing, and node lifecycle.
  • Own the stateful data plane: hosted vector stores for semantic memory and graph stores for knowledge graphs - deployed, backed up, scaled, and recovered, with restores that are tested rather than hoped for.
  • Oversee the sandbox runtime and its host-side control plane: lifecycle, exec, snapshot/fork, teardown, metering, and the threat model of the isolation boundary.
  • Direct the FastAPI control-plane services, Terraform/OpenTofu, Bicep, and the dashboards
  • Run the layer the backend services team builds on, expect to debug into their services, and expect them to read your dashboards.
  • Manage the open-source posture: build to OSS standards and release as it matures, with reviewed PRs, real docs, and reproducible builds.

Core Qualifications
  • 5+ years of shipping production systems in a systems language. Rust is the house language, but polyglots are welcome - deep Go, C/C++, or Zig with genuine appetite for Rust counts. Async runtimes, memory-safety discipline, and debugging at the syscall boundary should be familiar territory.
  • Operated Kubernetes workloads that other people depended on - controllers or operators, scheduling, autoscaling, node lifecycle. Youve been paged, and the experience changed how you build.
  • Strong ability to threat-model isolation boundaries (namespaces, cgroups, seccomp, hypervisors), including identifying what an untrusted guest could observe, forge, or exhaust, and applying security best practices for agentic execution - least privilege, no credentials in the sandbox, audit trails, and human approval on write actions.
  • Hands-on experience deploying or operating open-weight LLM serving infrastructure (vLLM/SGLang or similar), including packaging models into reliable, metered production endpoints.
  • Performance discipline in distributed systems: you measure before you optimize, and you can tell the story of a latency you killed with the numbers attached.
  • Practical depth in some of our core stack - Rust (tokio), Python/FastAPI, Kubernetes operators (controller-runtime/Kubebuilder/CRDs), KEDA, Karpenter, GPU device plugins/DRA - with a genuine willingness to research your way into the rest.
  • Familiarity with isolation technology (Firecracker, Kata, gVisor, or comparable), secrets management and egress control (Vault/KMS-class), and hosting stateful systems (vector stores like pgvector/Qdrant, graph stores like Neo4j) with backup and failover discipline.
  • Comfort operating across cloud and GPU substrates - AWS/Azure/GCP plus managed GPU clouds - using infrastructure-as-code (Terraform/OpenTofu, Bicep) and observability tooling (OpenTelemetry).
  • Bachelors Degree; Masters is a plus

Why AZX!
  • Be part of a fast-growing, profitable, mission-driven company with industry-leading clients tackling the massive opportunity of AI transformation in critical industries.
  • Competitive early-stage startup compensation (based on capabilities, experience, and location)
  • Bonus eligibility
  • Health insurance with meaningful coverage for dependents
  • Flexible paid time off
  • Equity
  • Fully remote culture with a cluster of teammates in Seattle


Additional Information:
  • Must be able to travel 2x/year for company summits
  • Applicants must be currently authorized to work in the United States on a full-time basis.
  • We are unable to sponsor or take over sponsorship of employment visas at this time.
  • Please note that our interview process includes a written take-home assignment followed by a live two-hour technical session with our engineering team, so if that format isnt a good fit, wed ask that you not apply
  • Please only apply to a maximum of 2 roles at a time, any applicants who apply to more then 2 roles within a 6 month period will automatically be disqualified


Next Steps:

If this job sounds like a great fit but you dont check ALL of these qualification boxes, wed still love to hear from you!

Similar Jobs

More Jobs at AZX

More Information Technology Jobs

Find similar Senior Software Engineer (AI Inference & Runtime Platform) jobs: