Senior Site Reliability Engineer

Luma

$150K — $180K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years as an SRE or infrastructure engineer in a fast-paced environment.
  • Deep expertise in Linux and low-level performance debugging.
  • Experience with Terraform, Airflow, and Ray.
  • Strong AWS or OCI experience.
  • Practical knowledge of high-performance networking (InfiniBand, RDMA).
  • Familiarity with security best practices and compliance frameworks.

Responsibilities

  • Manage GPU clusters for AI training and inference across AWS and OCI.
  • Participate in re-architecture sessions for system efficiency and scalability.
  • Tweak Linux performance at the OS and kernel level.
  • Develop automation scripts in Python, Go, or Bash to improve infrastructure management.
  • Act as the final escalation point for complex GPU and networking issues.
  • Ensure compliance with security standards like SOC 2 and ISO.

Benefits

  • Remote work flexibility within the US.
  • Opportunity to work with cutting-edge GPU technologies.
  • Engagement in high-impact architectural redesign projects.
  • Collaboration with industry vendors like NVIDIA.
  • Involvement in maintaining essential security certifications.
Full Job Description
Team: Infra Reliability • SF Bay Area / Remote (US)

You'll own the GPU infrastructure Luma's research and product run on - thousands of NVIDIA and AMD GPUs across on-prem and multi-cloud (AWS and OCI). As a Senior SRE, you keep training and inference clusters reliable and fast, and you help redesign them for the next level of scale.

This is a hands-on, close-to-the-metal role for a first-principles Linux engineer. You'll be the final escalation for the hardest GPU, networking, and kernel-level failures, sometimes debugging directly with NVIDIA. It fits someone who thrives on low-level problems in a fast, less-structured environment. If you want a narrow, well-bounded ops role, this isn't it.

What You'll Own
  • Take end-to-end ownership of production GPU clusters for training and inference across AWS and OCI, keeping them highly available and performant.
  • Join critical re-architecture sessions to redesign systems for higher efficiency and scale.
  • Tune Linux performance deeply, at the OS and kernel level.
  • Build automation in Python, Go, or Bash to manage, monitor, and self-heal infrastructure without heavy toil.
  • Serve as the final escalation for the hardest GPU, networking (InfiniBand/RDMA), and system failures, working with vendors like NVIDIA.
  • Help achieve and maintain security certifications (SOC 2 Type 1 & 2, ISO) with strong infrastructure security practices.

First 90 Days

One way the first 90 could unfold.
  • Days 1-30 - Immerse & Diagnose: Learn the current clusters across on-prem, AWS, and OCI, and where reliability and performance hurt most.
  • Days 30-60 - Ship & Validate: Take ownership of a production cluster and ship automation or tuning that measurably improves availability or performance.
  • Days 60-90 - Scale & Systemize: Contribute to the next-gen re-architecture and harden security and compliance practices.

What You Bring
  • 5+ years as an SRE, production, or infrastructure engineer in a fast-paced, large-scale environment.
  • Deep, hands-on Linux expertise, containerized systems, and low-level performance debugging.
  • Working experience with Terraform, Airflow, and Ray.
  • Strong experience with AWS or OCI.
  • Practical experience with high-performance networking (InfiniBand, RDMA, or RoCE).
  • Working knowledge of security best practices and compliance frameworks like SOC 2 and ISO.
  • Comfort in a less-structured, fast-paced environment.

Nice to Have
  • Deep expertise with GPU tooling for NVIDIA and AMD (DCGM, ROCm).
  • Experience managing large-scale GPU clusters for AI/ML training or inference.
  • Familiarity with Kubernetes or orchestration frameworks like Ray.
  • Deep expertise in data pipelines and infrastructure.

Similar Jobs

More Jobs at Luma

More Information Technology Jobs

Find similar Senior Site Reliability Engineer jobs: