BeyondTrust

Staff Site Reliability Engineer

BeyondTrust$135K — $160K *
US-Anywhere
+ 2 other locationsRemote
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years experience in SRE, DevOps, or Platform Engineering
  • At least 2 years in a Senior or Staff level role
  • Proven management of automation and infrastructure in cloud and on-premises
  • Experience with Docker and Kubernetes administration
  • Proficiency in a systems programming language (Go, Java, or C#)
  • Deep understanding of open telemetry and monitoring best practices
  • Familiarity with GitOps workflows and Infrastructure as Code best practices

Responsibilities

  • Design and maintain highly available and secure systems in cloud and on-premises environments
  • Champion initiatives that reduce cognitive load for software engineers
  • Own and optimize foundational services like api gateways and service meshes
  • Drive a culture of 'Everything as Code' with automated deployment workflows
  • Standardize and secure CI/CD pipelines for rapid code deployments
  • Implement chaos engineering frameworks for proactive system reliability
  • Architect and enhance telemetry and observability stacks for system visibility
  • Partner with leadership to define the long-term SRE strategy
  • Mentor engineers on reliability practices and improve documentation

Benefits

  • Opportunity to influence the engineering roadmap
  • Engagement in high-impact initiatives with a focus on AI-first operations
  • Access to advanced technologies and cutting-edge tools
  • Robust mentorship and professional development opportunities
  • Collaborative environment with cross-functional teams
Full Job Description
The Role

We are seeking a Staff Site Reliability Engineer (SRE) to lead the evolution of the Password Safe platform, infrastructure, and deployment ecosystem. In this role, you will bridge the gap between systems engineering and software development, architecting highly availability highly resilient systems across both cloud and on-premises environments.

As a Staff level engineer, you will tackle complex technical challenges and serve as a technical leader, mentoring engineers and influencing our broader engineering roadmap. You will own the reliability, scalability, and efficiency of shared services, CI/CD pipelines, and platform engineering initiatives via an AI first operating model.

What You'll Do

Infrastructure & Platform Engineering
  • Design, scale, and maintain highly available, secure, and resilient systems spanning both cloud (AWS/Azure) and on premises environments.
  • Champion platform engineering initiatives that reduce cognitive load for software engineers and accelerate velocity.
  • Own and optimize foundational services, including api gateways, service meshes, caches, configuration management, and secrets management.
  • Drive a "Everything as Code" culture, ensuring that all cloud and on-prem infrastructure, CI/CD pipelines, and configurations are declaratively defined, version-controlled, and deployed via automated GitOps workflows.

CI/CD & Automation
  • Standardize, secure, and optimize modern CI/CD pipelines to ensure safe, repeatable, and rapid code deployments.
  • Treat infrastructure as code (Terraform, OpenTofu, or Ansible), driving architectural patterns that eliminate configuration drift.

Reliability, Testing & Observability
  • Design and implement automated chaos engineering frameworks and disaster recovery simulations to proactively identify system weaknesses.
  • Architect and mature our telemetry stack (metrics, logs, traces) using tools like Grafana Cloud, Datadog, and OpenTelemetry to ensure deep system visibility.
  • Define, implement, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs) across critical applications.

Leadership
  • Partner with engineering leadership to define the long-term SRE strategy.
  • Raise the bar through mentorship and standards. Coach engineers on reliability practices, run design and incident reviews, and build documentation and tooling that makes reliability knowledge accessible.

What You'll Bring
  • Experience: 7+ years of experience in SRE, DevOps, or Platform Engineering, with at least 2 years in a Senior or Staff level.
  • Systems Architecture: Proven track record of managing automation and infrastructure across both cloud and on-premises environments.
  • Containerization: Experience with Docker and Kubernetes (including cluster administration, networking, and security primitives).
  • Release Orchestration: Prior experience with release orchestration strategies such Canary or Blue Green models.
  • Software Engineering: Proficiency in at least one systems language (e.g. Go, Java, or C#) to build automation, tooling, and internal APIs.
  • Observability: Deep understanding of open telemetry and monitoring and observability best practices (e.g. Golden vs Red Signals, and App monitoring).
  • Everything as Code: Familiarity with GitOps workflows, Infrastructure and Config as Code best practices.

Nice To Have
  • Knowledge of UI automation testing.
  • Experience designing and testing microservice based applications.
  • Experience working with virtual machines and managing test environments.
  • Experience working in a continuous integration environment.
  • Deep understanding of Linux and Windows internals.
  • Experience migrating workloads seamlessly across on-prem and cloud environments.
  • Understanding of modern DevSec Ops practices.

Better Together

About BeyondTrust

BeyondTrust is a cybersecurity company that provides solutions for privileged access management, vulnerability management, and remote access management. The company was founded in 2006 and is headquartered in San Diego, California. BeyondTrust's products are used by organizations in various industries, including finance, healthcare, and government. The company has received numerous awards for its products, including the SC Award for Best Vulnerability Management Solution and the CRN Tech Innovator Award for Privileged Access Management.
Learn more about BeyondTrust
Size
1,000 employees
Industry

Similar Jobs

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: