Member of Technical Staff - Infrastructure

Gimlet Labs

• $150K — $180K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in infrastructure, cluster engineering, or distributed systems
  • Deep understanding of Linux systems including performance debugging and kernel issues
  • Hands-on expertise with Kubernetes, Slurm, or similar orchestration tools
  • Strong automation skills using Terraform, Ansible, or Python
  • Familiarity with GPU/accelerator infrastructure and relevant software stacks
  • Experience with high-performance networking technologies like InfiniBand or high-speed Ethernet
  • Ability to navigate and thrive in a fast-paced startup environment

Responsibilities

  • Design, deploy, and operate large-scale clusters for AI inference
  • Automate provisioning, configuration, upgrades, and lifecycle management
  • Scale heterogeneous bare-metal provisioning systems across datacenters
  • Debug production issues in Linux, networking, storage, and orchestration layers
  • Build networking infrastructure with a focus on performance and RDMA
  • Develop observability metrics for cluster health and performance
  • Enhance reliability and recovery of multi-node production systems
  • Collaborate with teams to support high-throughput AI workloads

Benefits

  • Opportunity to work with cutting-edge AI infrastructure technology
  • Hands-on role in a dynamic startup environment
  • Collaborative culture with cross-functional teamwork
  • Professional growth in emerging areas of heterogeneous computing
  • Focus on operational excellence and scalable infrastructure practices
Full Job Description
About the role

As an Infrastructure Platform Engineer, you will build the systems that turn heterogeneous accelerator hardware into reliable production infrastructure for Gimlet's AI cloud.

Gimlet's fleet spans hardware with different architectures, software stacks, operational characteristics, and failure modes. Your work will determine how new hardware is brought online, how clusters are provisioned and operated, and how production inference systems remain reliable as the fleet scales.

You will work across bare metal, Linux, Kubernetes, cluster scheduling, observability, and automation. You will build systems that abstract differences between accelerator architectures, make new hardware production-ready, and improve the reliability and operability of Gimlet's infrastructure.

What success looks like

In your first 12-18 months, you will:
  • Deploy and operate production clusters across different accelerator architectures
  • Automate hardware provisioning, validation, upgrades, and fleet lifecycle management
  • Improve cluster scheduling, resource utilization, isolation, and capacity management
  • Build observable infrastructure that enables faster debugging, incident response, and recovery
  • Partner across distributed systems, runtime, compiler, networking, and hardware teams to bring new accelerators into production

You may be a good fit if you have
  • Experience in infrastructure, cluster engineering, platform engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Experience automating infrastructure with Python, Go, Terraform, Ansible, or similar tools
  • Experience with GPU or accelerator infrastructure, including drivers, firmware, or CUDA/ROCm
  • The ability to build systems that are observable, recoverable, and reliable in production
  • A bachelor's degree in a relevant field or equivalent practical experience

Strong candidates may also have
  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure
  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up
  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting
  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks
  • Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling
  • Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators

Agency Policy: Gimlet Labs does not accept unsolicited resumes from recruitment agencies or search firms. Any unsolicited resumes submitted without a signed agreement will be considered the property of Gimlet Labs, and no fees will be paid.

Similar Jobs

More Jobs at Gimlet Labs

More Enterprise Technology Jobs

Find similar Member of Technical Staff - Infrastructure jobs: