Head of AI Infrastructure

Lawrence Harvey

• $250K — $350K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years in infrastructure with hands-on experience in physical production compute at scale
  • Experience at a hyperscaler, GPU cloud, or neocloud
  • Practical AI inference knowledge, particularly vLLM or serving frameworks
  • Depth in inference serving, InfiniBand or RoCE fabrics, distributed storage, or bare-metal Kubernetes
  • Experience running distributed or multi-site infrastructure
  • Strong Linux skills, plus proficiency in Python or Go
  • Technical leadership experience and desire to build and manage a team

Responsibilities

  • Bring up, burn in, and run GPU pods in production, managing acceptance benchmarks
  • Measure and enhance inference performance across various metrics
  • Test and roll out updates safely with clear benchmarks and rollback plans
  • Operate multi-tenant, bare-metal Kubernetes GPU platform with 24/7 incident response
  • Design infrastructure for graceful failure and coordination between pods
  • Write operational runbooks for field technicians at unmanned sites
  • Act as technical lead for customers and define the 12-18 month infrastructure roadmap
  • Build and lead the infrastructure team as fleet scales

Benefits

  • Fully remote work option within the US or Canada
  • Opportunity for travel to pod sites and customer locations
  • Participation in two company-wide gatherings per year
  • Equity options as part of the compensation package
Full Job Description
Head of AI Infrastructure

United States

$250-350K + Bonus + Equity

Must be a citizen/authorized to work in the US/Canada

Remote with travel

The role

This is the company's first dedicated infrastructure hire, reporting directly to the CTO.

You'll start hands-on and own the AI platform end to end: compute, memory, network, storage, orchestration and field operations. As the fleet grows, you'll build and lead the infrastructure team. You'll also be the senior technical voice for customers and set the 12 to 18 month infrastructure roadmap.

The work moves fast. New NVIDIA drivers ship weekly, new vLLM versions every two weeks and new models monthly, so continuous benchmarking and safe rollouts are at the heart of the job.

What you'll do
  • Bring up, burn in and run GPU pods in production, and own the acceptance benchmarks every new pod and hardware generation must pass
  • Measure and improve inference performance across compute, memory, network and storage: tokens per second per kW, KV-cache offload, RoCEv2 and NCCL fabric performance, model cold-start
  • Test and roll out frequent driver, vLLM and model updates safely, with clear benchmarks and rollback plans
  • Run power-aware operations: GPU power caps, curtailment and workload drain coordinated with other on-site energy loads
  • Design for graceful failure, so work hands off cleanly between pods and sites
  • Operate a multi-tenant, bare-metal Kubernetes GPU platform against service level objectives, with 24/7 incident response
  • Write the runbooks field technicians follow at unmanned sites
  • Act as technical lead for customers, and set the 12 to 18 month infrastructure roadmap
  • Hire, build and lead the infrastructure team as the fleet scales

What you'll bring
  • 12+ years in infrastructure, with hands-on ownership of physical production compute at scale (not only consuming public cloud services)
  • Experience at a hyperscaler, GPU cloud or neocloud: you know how large organizations keep large systems running
  • Practical AI inference knowledge, such as vLLM or other model serving frameworks
  • Depth in at least two of: inference serving, InfiniBand or RoCE fabrics, distributed storage, bare-metal Kubernetes
  • Experience running distributed or multi-site infrastructure
  • Strong Linux skills, plus Python or Go: you read the code and write the fix
  • Technical leadership experience, and the ambition to build and manage a team

Nice to have
  • Edge or distributed-site infrastructure, such as CDN points of presence, cloud local zones or telecom edge
  • Experience with 1,000+ GPU clusters
  • Power-aware scheduling, demand response or GPU cluster power management
  • Background at a GPU, chip or infrastructure vendor, such as NVIDIA, AMD, Intel or Oracle

You don't need to check every box to be considered.

Location, compensation and how to apply
  • Location: fully remote, anywhere in the US or Canada
  • Travel: to pod sites and customers as needed, plus two company-wide gatherings a year
  • Work authorization: must already be authorized to work in the US or Canada; visa sponsorship is not available
  • Compensation: competitive base salary with significant equity

Similar Jobs

More Jobs at Lawrence Harvey

More Information Technology Jobs

Find similar Head of AI Infrastructure jobs: