Software Engineer, High Performance Computing

Eventual Computing

$130K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in systems engineering or related field
  • Strong programming skills in Rust, C++, or C
  • Deep familiarity with operating systems and memory hierarchies
  • Expertise with NVMe, networking, CPU, and memory interactions
  • Ability to optimize I/O performance and manage data paths effectively

Responsibilities

  • Design and build a high-performance video-native dataloader for GPUs
  • Profile and optimize the data flow across various storage and memory systems
  • Maximize utilization of cutting-edge hardware during training tasks
  • Establish performance benchmarks in collaboration with research teams
  • Integrate solutions with partner labs to measure performance end-to-end
  • Collaborate with cross-functional teams on data indexing and model ingestion

Benefits

  • In-person work with a close-knit team in SF 4 days a week
  • Competitive compensation and startup equity
  • Catered meals for on-site employees
  • Commuter benefits available
  • Team-building events and recreational activities
  • Comprehensive health, vision, and dental insurance
  • Flexible paid time off
  • Access to the latest Apple technology
  • 401(k) plan with employer matching
Full Job Description
Your Role

As a Systems Engineer on the Dataloading team, you'll build the layer that turns multi-petabyte video corpora into dict[str, Tensor] already on the GPU at line rate. We work with the top labs training Physical AI on the newest generation hardware - H100, B200, GB200, NVL72, with Vera Rubin on the horizon - on billions of dollars worth of compute, in collaboration with partners that are the largest public AI companies on Earth. Our job is to keep those GPUs fed: rank-aware sampling, NVMe caching, video and sensor co-loading, random access into clips, decode pipelining. Streaming alone can already saturate a B200; the hard part is enabling the complex sampling patterns researchers actually need without giving up a single percentage point of MFU.

This is a systems engineering role for someone who feels physical pain when a system is slow. You won't need GPU experience on day one - we'll uplevel you on NVL72, CUDA, and SLURM. We will need you to bring real expertise on what happens between NVMe, network, memory, and CPU, and a deep instinct for where bytes go.

Key Responsibilities
  • Design and build the video-native dataloader: rank-aware, NVMe-cached, random-access into clips, returns tensors directly to the GPU.
  • Profile and optimize the full data path from object store 14 NVMe 14 page cache 14 host RAM 14 device RAM. Eliminate every avoidable copy and stall.
  • Saturate the latest hardware (B200, GB200, NVL72) on real customer training jobs. Push toward Vera Rubin bandwidth requirements.
  • Own performance benchmarks against customer baselines (custom DataLoaders, DALI, decord, LeRobot) and against our own historical numbers - regressions get caught at PR time.
  • Partner with researchers at our partner labs to land the loader in their training stack and measure MFU end-to-end.
  • Work cross-team with Storage Infrastructure on the index/format boundary and with Visual Understanding on the model-output ingestion path.


What we look for
  • Obsession with systems-level performance. You can recite Jeff Dean's 2numbers every programmer should know2 in your sleep. You eat flamegraphs for breakfast.
  • Strong opinions on io_uring - love it or hate it, you've earned the opinion.
  • Live and breathe Rust, C++, or C. You reach for them when it matters and you know why.
  • Strong familiarity with operating systems - page cache, scheduling, syscalls, NUMA, memory hierarchies.
  • A sense for where bytes actually go: NVMe vs. memory vs. network vs. PCIe vs. NVLink, and the throughput and latency budgets of each.


Nice to have
  • Experience working with GPUs is a plus, but you don't need it on day one.
  • Experience working with SLURM, Kubernetes for GPU workloads, or other HPC schedulers.
  • Hands-on CUDA experience.
  • Deep expertise on memory and caching subsystems - page cache tuning, hugepages, NUMA pinning, GPU-Direct Storage.
  • Worked on video decode pipelines (PyAV, decord, NVDEC) or PyTorch DataLoader internals.
  • Contributed to open-source systems projects in Rust/C++.


Perks & Benefits
  • In-person, tight-knit team - 4 days/week in our SF Mission office.
  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team-building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment.
  • 401(k) plan with match.


If slow systems evoke emotional pain for you and you want to spend the next few years making the most expensive GPU clusters on the planet earn their keep, we'd love to talk.

Similar Jobs

More Jobs at Eventual Computing

More Information Technology Jobs

Find similar Software Engineer, High Performance Computing jobs: