Member of Technical Staff - Compute Cluster

Causal Labs

$130K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm)
  • Strong systems background with Linux, networking, storage, and infrastructure-as-code expertise
  • Familiarity with cloud platforms (GCP, AWS, or Azure) focusing on their ML/AI services
  • Proficient in monitoring, logging, observability, and version control best practices for ML systems
  • Knowledgeable in CUDA/NCCL and performance profiling techniques for distributed workloads
  • Demonstrated ability to take ownership of deliverables from requirements through execution.

Responsibilities

  • Design, deploy, and operate large distributed GPU clusters from provisioning to upgrades
  • Extend scheduling and orchestration systems for optimized resource management
  • Develop software that simplifies cluster management for researchers and engineers
  • Manage cluster storage and artifact paths, ensuring clear retention and lineage
  • Enhance reliability and error recovery processes; implement observability measures
  • Collaborate with researchers to facilitate large-scale runs and optimize performance

Benefits

  • Flexible work environment
  • Opportunities for professional growth and development
  • Access to cutting-edge technology and resources
  • Collaborative culture with passionate colleagues
  • Impactful work that drives innovation in research
Full Job Description
We look for infrastructure engineers who are excited to tackle unsolved problems. Everything we do - training, evaluation, serving - runs on our GPU fleet. Your mission is to design, build, and operate the supercomputing environment underneath it all, delivering performant, reliable, and cost-efficient compute to ensure research is able to iterate rapidly at scale.

Responsibilities
  • Design, deploy, and operate large distributed GPU clusters end to end: provisioning, imaging, upgrades, and capacity planning
  • Extend scheduling and orchestration systems (e.g. Kubernetes, Slurm) for topology-aware placement, preemption, quotas, and multi-tenancy across training and inference workloads
  • Build software that abstracts cluster management and presents a unified, self-serve interface to researchers and engineers
  • Own cluster storage and artifact paths for checkpoints and logs, with clear retention and lineage
  • Monitor and continuously improve reliability and error recovery; build the observability to catch failures before researchers do
  • Partner with researchers to unblock large-scale runs and advise on performance and placement trade-offs


What we're looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
  • Experience operating large-scale GPU clusters and container orchestration frameworks (e.g. Kubernetes, Slurm, Docker)
  • Strong systems background: Linux, networking, storage, infrastructure-as-code
  • Knowledge of cloud platforms (GCP, AWS, or Azure) and their ML/AI service offerings
  • Understanding of monitoring, logging, observability, and version control best practices for ML systems
  • Familiarity with CUDA/NCCL and performance profiling for distributed workloads
  • Owns deliverables end-to-end, from requirements through autonomous execution

Similar Jobs

More Jobs at Causal Labs

More Information Technology Jobs

Find similar Member of Technical Staff - Compute Cluster jobs: