Site Reliability Engineer II

Varda Space Industries

$133K — $170K *
Aerospace & Defense
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's in computer science, engineering, or STEM field with 3+ years in SRE, or 5+ years in DevOps/SRE without degree.
  • Hands-on experience with Infrastructure as Code (IaC) using Terraform.
  • Proficient with Kubernetes or similar container orchestration in production.
  • Familiar with observability tools like Prometheus and Grafana.
  • Scripting skills in Python, Bash, or PowerShell.
  • Strong written and oral communication skills.

Responsibilities

  • Deploy and manage critical applications and infrastructure for spacecraft and company operations.
  • Develop Infrastructure as Code (IaC) frameworks with Terraform.
  • Implement observability systems including metrics and alerting.
  • Create and maintain CI/CD pipelines for efficient deployment.
  • Collaborate with engineers to ensure reliable and scalable systems.
  • Analyze and fix system bottlenecks and risks, enhancing performance and stability.
  • Address production incidents and conduct root cause analyses.

Benefits

  • Flexible PTO policy plus 12 paid holidays.
  • 100% company-paid medical, dental, and vision insurance for employees and dependents.
  • 12 weeks of parental leave and family support through Maven Clinic.
  • 401(k) with 6% employer match and immediate vesting.
  • Daily lunch and regular team bonding events.
Full Job Description
About This Role

At Varda Space Industries, we're pushing the boundaries of what's possible in space and materials science - and we're looking for bold engineers to help us get there. As a Site Reliability Engineer, you'll be critical in building, scaling, and maintaining the infrastructure that powers our systems on Earth, in orbit, and everything in between.

We are looking for an experienced engineer with deep working knowledge of Kubernetes and containerized technologies. You are a hands-on operator and builder who applies first-principles thinking to both software delivery (DevOps) and production reliability (SRE), and thrives in complex, mission-critical environments.

In this role, you will:
  • Solve challenging technical problems across a wide range of modern technologies.
  • Apply a software engineering mindset to automate operations and improve system reliability, scalability, and resilience
  • Design and build infrastructure that enables rapid development - from cloud-based services to embedded software running on spacecraft.
  • Shape Varda's infrastructure strategy and drive operational excellence across containerized and modernized environments.

Responsibilities
  • Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems.
  • Build and evolve Infrastructure as Code (IaC) frameworks using tools such as Terraform
  • Implement and operate observability systems (metrics, logging, tracing) and actionable alerting.
  • Build and maintain CI/CD pipelines to enable safe, repeatable, and rapid deployments.
  • Partner with software and hardware engineers to deliver highly operable, reliable, and scalable systems and pipelines; ensuring they have the tools and infrastructure needed for rapid iteration.
  • Identify, analyze, and resolve system bottlenecks and reliability risks; perform performance tuning and implement long-term stability improvements.
  • Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems.
  • Rotate through the team's on-call schedule to keep critical systems healthy and responsive.
  • Must be willing to work extended hours and weekends as needed
  • Occasionally travel to customer sites and other Varda locations to troubleshoot, deploy, or test critical infrastructure.

Basic Qualifications
  • Bachelor's degree in computer science, engineering, or related STEM field with 3+ years of Site Reliability Engineering experience, or 5+ years of progressive experience in DevOps, SRE, or Systems Engineering in lieu of a degree.
  • Experience with Infrastructure as Code (IaC) using tools like Terraform to automate server provisioning
  • Experience operating Kubernetes or similar container orchestration platforms in production environments.
  • Experience with Prometheus, Grafana, InfluxDB, or similar technologies.
  • Knowledge of software-defined networking (VPC, Subnets, Firewalls, VPNs, etc.)
  • Python, Bash, PowerShell (or similar) scripting experience
  • Positive and strong communication skills, both written and oral

Preferred Skills and Experience
  • Experience provisioning and managing scalable Azure cloud infrastructure using native tools and best practices
  • Experience implementing configuration management, provisioning, and workflow automation solutions via Infrastructure as Code, CI/CD, and GitOps (e.g., Ansible, Salt, Argo CD, etc.).
  • Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker)
  • Experience with GPU workloads or high-throughput computing.
  • Hands-on experience operating and optimizing High-Performance Computing (HPC) environments, including workload schedulers such as Slurm (e.g., queue/partition design, fair-share scheduling, and cluster resource management).
  • Experience with hybrid environments (cloud + on-prem or edge systems)
  • Experience debugging distributed systems at scale (network, storage, latency)
  • Experience with databases and data modeling

Additional Requirements

Must be physically able to regularly lift 25 lbs. for duties such as delivering computers, unpacking and rack-mounting equipment, etc.
Pay Range
  • Site Reliability Engineer: 133,000 - $170,00 per year
  • This role is on-sitein El Segundo, CA
  • Leveling and base salary are determined by job-related skills, education level, experience level, and job performance
  • You will be eligible for long-term incentives in the form of stock options and/or long-term cash awards


Benefits

Varda offers a comprehensive benefits package designed to support health, financial well-being, and a high-quality workplace experience. Below is an overview of what full-time employees receive (at this time, interns receive a subset of benefits):

Health & Wellness
  • Flexible PTO policy + 12 paid holidays
  • 100% company-paid Medical, Dental, and Vision insurance plans for employees and dependents with FSA and employer-matched HSA options
  • Voluntary accident, hospital, critical illness, and pet insurance
  • $120/month wellness reimbursement for gym and fitness expenses
  • 12 weeks of parental leave (with supplemental disability leave for CA mothers)
  • Family building, pregnancy, parenting and menopause benefits via Maven Clinic
  • Sponsored One Medical memberships for employees and their dependents

Financial & Retirement
  • Substantial incentive equity in a fully funded space start-up
  • 401(k) retirement plan with 6% employer match (immediately vested)
  • $20/pay period cell phone reimbursement
  • Relocation support for new hires, if needed

Workplace Experience & Perks
  • Fully stocked kitchen with lunch provided daily and dinner provided twice weekly
  • Company and team-bonding events, happy hours and mission-success celebrations
  • Complimentary EV charging


Similar Jobs

More Jobs at Varda Space Industries

More Aerospace & Defense Jobs

Find similar Site Reliability Engineer II jobs: