Umbra

Senior Site Reliability Engineer

Umbra$150K — $180K *
Aerospace & Defense
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science or a related technical field.
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform.
  • Expertise in managing distributed systems and AWS services like EC2, S3, and Lambda.
  • Proficiency in Kubernetes cluster management in production environments.
  • Experience using Terraform and Infrastructure-as-code (IaC) practices.
  • Leadership experience in Agile/Scrum environments with proven project success.
  • Skills in developing infrastructure monitoring and alerting strategies.

Responsibilities

  • Ensure reliability and scalability of critical systems through proactive monitoring and incident response.
  • Develop and promote new technologies and tools that enhance team capabilities.
  • Foster a culture of excellence and reliability within the team.
  • Continuously improve team processes and workflows for efficiency.
  • Collaborate with cross-functional teams and stakeholders to align on technical strategies.
  • Participate in on-call rotations to support and resolve technical issues.

Benefits

  • Flexible Time Off, Sick, Family & Medical Leave
  • Comprehensive medical, dental, vision, and life insurance (employer funded)
  • 401k with a 3% non-elective company contribution
  • Stock options and opportunities for employee investments
  • Free parking and daily lunch provided in the office
Full Job Description
About the Job

We are seeking an experienced SeniorSite Reliability Engineer to help design, build, operate, and scale the mission- and business-critical infrastructure that powers Umbra's systems. In this role, you will leverage a deep understanding of modern infrastructure, distributed systems, and the broader technology stack to drive technical excellence, make thoughtful architectural decisions, and balance long-term scalability with operational reliability.

You'll partner closely with engineering teams to improve processes, champion best practices, evaluate emerging technologies, and implement solutions that enhance the performance, resilience, and efficiency of our platforms. The ideal candidate is a collaborative technical leader who communicates effectively across technical and non-technical teams and drives meaningful improvements that have a lasting impact across the organization.

This position is based on-site in either our Arlington, VA office, Reston, VA office or Santa Barbara/Goleta, CA office.

Key Responsibilities
  • Ensure the reliability and scalability of critical systems, meeting SLAs through proactive monitoring and effective incident response.
  • Develop and promote new technologies and tools, conducting research and creating proofs of concept to introduce solutions that enhance the team's capabilities.
  • Lead by example in fostering a culture of excellence and reliability.
  • Continuously evaluate and improve team processes and workflows to increase efficiency and reduce complexity.
  • Collaborate closely with cross-functional teams, product managers, and stakeholders to align on technical strategy and provide expert guidance.
  • Participate in on-call rotations, providing support and resolving complex technical issues.

Requirements
Required Qualifications
  • Bachelor's degree in Computer Science or a related technical field.
  • 5-8+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems.
  • Extensive experience with AWS services (EC2, S3, Lambda, VPC Networking) and deep knowledge of cloud infrastructure, networking, and security best practices.
  • Proficiency running, optimizing, and scaling Kubernetes clusters in production environments.
  • Experience using and writing Terraform to architect and manage production infrastructure.
  • Ability to create and utilize Infrastructure-as-code (IaC), GitOps practices, and automation tools to increase reliability and reduce manual tasks.
  • Proven success in leading teams or projects using Agile/Scrum methodologies.
  • Expertise in infrastructure and software architecture, capable of designing and implementing large-scale, reliable systems with minimal guidance.
  • Experience developing and managing comprehensive infrastructure monitoring and alerting strategies.
Desired Qualifications
  • 10+ years in a Site Reliability Engineer or DevOps role supporting a SaaS platform, with demonstrated expertise managing distributed systems.
  • Advanced understanding of cloud and application security, identity management, and compliance.
  • Expertise in service mesh and service registration technologies, focusing on performance and reliability.
  • Experience in the aerospace industry.


Benefits
  • Flexible Time Off, Sick, Family & Medical Leave
  • Medical, Dental, Vision, Life, LTD, STD (employer funded)
  • Vol Life, Critical Illness, Accidental, Hospital Indemnity, Pet Insurance (employee funded)
  • 401k with 3% non-elective company contribution
  • Stock Options
  • Free Parking
  • Free lunch daily in office


Pay Transparency
This job posting may cover multiple career levels. To ensure greater transparency, we provide base salary ranges for all roles, regardless of location. Our standard pay ranges are based on the role's function and level, benchmarked against similar growth-stage companies. Compensation may vary based on geographical location, as certain regions may have different cost-of-living factors. The final offer will also be influenced by the candidate's skills, responsibilities, and relevant experience.

Compensation Range

The Compensation Range for this role is $150,000 - $180,000 DOE.

About Umbra

Umbra is a computer hardware company that specializes in developing high-performance rendering software and hardware for the gaming and entertainment industries. The company was founded in 2006 and is headquartered in Ottawa, Canada. Umbra's technology is used by some of the world's leading game developers and studios to create immersive and realistic gaming experiences. The company's products include software tools for real-time rendering, as well as hardware solutions for high-performance graphics processing. Umbra's mission is to help game developers and studios create the most realistic and immersive gaming experiences possible.
Learn more about Umbra
Size
50 employees
Industry
Founded
2006

Similar Jobs

More Jobs at Umbra

More Aerospace & Defense Jobs

Find similar Senior Site Reliability Engineer jobs: