Staff Site Reliability Engineer

Finalsite

$120K — $145K *
US-Anywhere
+ 2 other locationsRemote
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in site reliability engineering or related fields
  • Expertise in Google Cloud Platform (GCP) and multi-cloud environments
  • Strong knowledge of Kubernetes and container platforms
  • Experience with Infrastructure as Code (IaC) tools like Terraform
  • Proven capability in incident management and leading response teams
  • Familiarity with observability practices and monitoring tools
  • Understanding of network architecture and security, especially with Cloudflare

Responsibilities

  • Guide the cloud architecture strategy focusing on scalability and cost-efficiency
  • Own the health and reliability of the Kubernetes platform
  • Design and evolve network architecture, including global routing
  • Develop Infrastructure as Code (IaC) standards for team adoption
  • Lead disaster recovery planning for platform reliability
  • Build tools that enable other teams to work independently
  • Represent infrastructure needs in collaborative planning discussions

Benefits

  • 100% remote work within the US
  • Opportunities for mentorship and growth in technical skills
  • Supportive team environment focused on collaboration
  • Access to cutting-edge technologies and cloud platforms
  • Chance to shape and influence best practices across the organization
Full Job Description
Job Description

SUMMARY

As a Staff Site Reliability Engineer, you'll focus on the platform behind Finalsite's Composer CMS, while also working cross-team to shape platform-wide standards and integrations. You'll help set the technical direction for how we build, scale, and operate our platform in a GCP-primary, multi-cloud environment. You'll be the person other engineers come to when a system needs to be rethought, not just repaired, and you'll play a key role in growing the senior engineers around you. This is a role for someone who wants to know exactly how the systems work, how they will fail, and how to build the guardrails that keep everyone else from finding out the hard way.
LOCATION

100% Remote - Anywhere within the US
WHAT WE LOOK FOR
  • Cloud Architecture & Scalability. You guide our cloud architecture strategy with a focus on scalability, maintainability, and cost efficiency, and you lead capacity planning so we scale ahead of demand instead of reacting to it.
  • Kubernetes & Container Platform Ownership. You own the health of our Kubernetes (GKE) platform, from cluster architecture to workload reliability at scale.
  • Network & Edge Architecture. You design and evolve our network architecture, including global routing, connectivity, and edge strategy with Cloudflare.
  • Infrastructure as Code. You drive IaC standards across the team, building reusable Terraform/Terragrunt modules that other teams can adopt without reinventing them.
  • Observability Mindset. You build and mature our observability practice, defining what we monitor, how we alert, and how we define and hold ourselves to SLOs.
  • Focus on Reliability. You lead disaster recovery planning for the platforms you own, including backup design, failover planning, and clear recovery objectives (RTO/RPO), and you architect highly available, fault-tolerant systems.
  • Security-First Engineering. You bring a security-first mindset to everything you build, treating it as a design input from day one, not a review gate at the end.
  • Cost Optimization. You keep a close eye on cloud cost and help teams make smart tradeoffs between performance, resilience, and spend.
  • Code Literacy. You can read, understand, and write basic code when needed, across languages and runtimes such as Ruby/Rails, Java, Python, for example. You're comfortable in application code to spot reliability, performance, or architecture issues.
HOW YOU'LL WORK
  • Team First. You mentor senior engineers, helping them grow their technical judgment and take on bigger calls of their own.
  • Developer Enablement. You build tools and patterns that let other teams move independently, turning one-off solutions into lasting, self-serve practice, so SRE isn't the bottleneck standing between people and the answer.
  • Collaborative Planning. You represent infrastructure and reliability concerns in planning conversations, weighing in early enough to shape decisions, not just implement them.
  • Change Management. You champion strong change management practices, including peer review, staged rollouts, and go/no-go gates for high-risk changes.
  • Incident Response. You lead response for high-severity incidents, participate in our on-call rotation, and turn every incident into a lasting improvement.
WE'D LIKE YOU TO HAVE
  • GCP & Infrastructure Expertise. Expert-level knowledge of GCP and infrastructure broadly, comfortable operating across the full stack rather than one layer of it.
  • Infrastructure as Code. Expertise in IaC, including Terraform and Terragrunt.
  • CI/CD Pipelines. Experience with GitLab (or similar) and CI/CD pipeline design and operation.
  • Kubernetes & Containerization. Expertise in Kubernetes and containerization, including Helm.
  • Network Architecture. Strong experience with network architecture, including routing and connectivity within GCP and across cloud providers.
  • Edge Security. Strong experience with edge security and traffic management, using solutions like Cloudflare.
  • Observability. Deep familiarity with observability tooling and monitoring strategy.
  • Incident Leadership. Real experience leading incident management and response, not just participating in it.
  • Staff-Level L6 Track Record. A track record of operating at a staff level, including mentoring senior engineers.
  • High Availability & Fault Tolerance. Experience architecting highly available, fault-tolerant systems.
  • Cost Allocation (Preferred). Experience with cloud cost allocation and control.
  • Capacity Planning (Preferred) . Experience with capacity planning and forecasting for infrastructure at scale.
  • Multi-Cloud Familiarity (Preferred). Familiarity with AWS and/or Azure (GCP-primary).
RESIDENCY REQUIREMENT

Finalsite offers 100% fully remote employment opportunities, however, these opportunities are limited to permanent residents of the United States. Current residency, as well as continued residency, within the United States is required to obtain (and retain) employment with Finalsite.

Similar Jobs

More Jobs at Finalsite

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: