Senior/Staff Kubernetes Infrastructure Engineer

Fal

$150K — $180K *
US-AnywhereRemote in United States
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience building and operating production Linux infrastructure
  • Strong production experience with Kubernetes on bare metal
  • Experience with Linux virtualization: KVM/QEMU, libvirt
  • Experience operating NVIDIA GPUs on Linux and Kubernetes
  • Strong networking fundamentals: TCP/IP, L2/L3
  • Practical scripting experience
  • Experience with configuration-management tools such as Ansible

Responsibilities

  • Design, automate, validate, and deliver complete lifecycle of customer compute environments
  • Use AI to automate and accelerate infrastructure delivery and operations
  • Provision dedicated Kubernetes and Slurm clusters for customer workloads
  • Build and maintain Linux images and automated OS-provisioning workflows
  • Operate the NVIDIA GPU stack for performance monitoring
  • Design Kubernetes and data-center networking using advanced technologies
  • Configure distributed and shared storage for high-performance workloads
  • Develop monitoring, alerting, diagnostics, and automated recovery solutions
  • Collaborate with customers to translate workload requirements into infrastructure designs

Benefits

  • Collaborative work environment
  • Opportunity to work with cutting-edge technology
  • Engagement in innovative projects
  • Professional development and growth opportunities
Full Job Description
You will build the high-performance compute environments we deliver to customers. These environments span bare-metal servers, virtual machines with GPU passthrough, Kubernetes and Slurm clusters, distributed storage, and high-speed networking.

You will work across the infrastructure stack - from Linux images and hardware provisioning to cluster networking, GPU performance, observability, and lifecycle automation. The goal is to make every customer environment performant, reliable, isolated, and repeatable.

Key Responsibilities:
  • Design, automate, validate, and deliver the complete lifecycle of customer compute environments - from provisioning through upgrades, recovery, and decommissioning
  • Use AI aggressively to automate and accelerate every aspect of infrastructure delivery and operations
  • Provision dedicated Kubernetes and Slurm clusters tailored to customer workloads
  • Build and maintain Linux images and automated OS-provisioning workflows
  • Operate the NVIDIA GPU stack: drivers, GPU Operator, NVIDIA Container Toolkit, device plugins, MIG, and GPU monitoring
  • Design Kubernetes and data-center networking using Cilium/Calico, MetalLB, VLAN, VXLAN, BGP, and ECMP
  • Configure distributed and shared storage for high-performance workloads
  • Build monitoring, alerting, diagnostics, and automated recovery for customer environments
  • Develop reusable tooling, standards, documentation, and runbooks
  • Collaborate with customers and internal teams to translate workload requirements into sound infrastructure designs

Requirements:
  • 5+ years of experience building and operating production Linux infrastructure
  • Strong production experience with Kubernetes on bare metal (bootstrapping, upgrades, HA control planes, etcd, containerd, CNI, CSI, ingress, load-balancing, observability, security, troubleshooting)
  • Experience with Linux virtualization: KVM/QEMU, libvirt, VFIO device passthrough
  • Experience operating NVIDIA GPUs on Linux and Kubernetes (drivers, container runtimes, device plugins, GPU Operator, GPU telemetry)
  • Strong networking fundamentals: TCP/IP, L2/L3, VLANs, routing, packet-level troubleshooting (tcpdump, Wireshark)
  • Practical scripting experience
  • Experience with configuration-management tools such as Ansible
  • Ability to diagnose complex, cross-layer infrastructure issues
  • Strong communication and ability to drive technical decisions across teams
  • Track record of moving quickly, taking ownership, and continuously improving systems

Nice to Have:
  • Production Slurm experience
  • High-performance networking: NVLink/NVSwitch, InfiniBand, RoCEv2, GPUDirect RDMA, NCCL, IMEX
  • Hugepages, NUMA, CPU pinning
  • SR-IOV, DPDK
  • Distributed storage: Ceph, Lustre, Weka
  • KubeVirt, OpenStack
  • IPsec, WireGuard, Tailscale
  • VXLAN, BGP, ECMP
  • Bare-metal management: BMC, IPMI, Redfish, PXE/iPXE, Kickstart, cloud-init
  • Network automation: NetBox, Nautobot, Nornir
  • AI training, inference, or distributed GPU workload infrastructure
  • Python or Go proficiency

Similar Jobs

More Information Technology Jobs

Find similar Senior/Staff Kubernetes Infrastructure Engineer jobs: