Staff DevOps Engineer

Cast & Crew

$190K — $235K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years of DevOps or platform engineering experience, with at least 2 at a Staff or Principal level.
  • Deep expertise in Kubernetes, specifically EKS, including troubleshooting and cluster operations.
  • Strong command of Azure DevOps Pipelines, with experience in YAML and environment promotion.
  • Proven experience in designing CI/CD systems for microservice architectures.
  • Experience operating observability platforms like New Relic or Datadog to enhance reliability.
  • Proficiency in a scripting language (Python, Bash, or Go) and Infrastructure-as-Code (Terraform or similar).
  • Excellent written communication skills for documentation and guidance.

Responsibilities

  • Architect and improve CI/CD pipelines in Azure DevOps.
  • Own AWS EKS cluster management, ensuring health and lifecycle.
  • Design Infrastructure-as-Code practices with Terraform, promoting GitOps.
  • Drive enhancements based on observability data, collaborating with SRE.
  • Define templates for containerized workloads and onboarding processes.
  • Act as escalation for complex infrastructure incidents, leading on-call and reviews.
  • Maintain and improve platform documentation and runbooks in Confluence.

Benefits

  • Comprehensive medical, dental, and vision coverage.
  • 401(k) match for financial growth.
  • Generous PTO and paid parental leave.
  • Health and wellness programs to support employee well-being.
  • Tuition reimbursement for continuing education.
  • Employee discounts for various services and products.
Full Job Description
Position Overview

We are looking for a Staff DevOps Engineer to serve as a technical anchor for our platform engineering practice. In this role you will own the design and evolution of our CI/CD pipelines, Kubernetes infrastructure on AWS EKS, and the developer experience tooling that hundreds of engineers depend on daily. Staff-level engineers at this organization are expected to operate with significant autonomy, identify and resolve systemic problems before they become incidents, and raise the technical bar across the teams they partner with.

Core Responsibilities

Platform & Infrastructure
  • Architect and continuously improve CI/CD pipelines in Azure DevOps, including pipeline-as-code standards, templating strategies, and artifact promotion workflows across environments.
  • Own the health and evolution of our AWS EKS clusters - node lifecycle, autoscaling, networking (VPC/CNI), RBAC, and cluster upgrades with minimal service disruption.
  • Design and enforce Infrastructure-as-Code practices using Terraform or equivalent tooling; champion GitOps patterns across engineering teams.
  • Drive platform reliability improvements informed by observability data from New Relic, working closely with SRE to translate dashboards and alerts into actionable platform changes.


Developer Experience
  • Define and maintain golden-path templates for containerized workloads - Dockerfile standards, Helm chart libraries, and local development parity with production.
  • Partner with engineering teams to accelerate onboarding of new services onto the platform and reduce toil through automation.


Incident & Operational Excellence
  • Act as an escalation point for complex infrastructure incidents coordinated through PagerDuty; participate in on-call rotation and lead post-incident reviews for platform-layer failures.
  • Identify recurring failure modes and drive systemic fixes that reduce page volume and MTTR across the platform.
  • Maintain and improve runbooks and platform documentation in Confluence, ensuring knowledge is accessible and current.


Technical Leadership
  • Define and socialize DevOps standards - pipeline design, container hygiene, secret management, and deployment safety - across a multi-team engineering organization.
  • Conduct architecture reviews and provide technical guidance on infrastructure-impacting decisions made by product engineering teams.
  • Mentor senior and mid-level engineers; grow internal platform capability through pairing, code review, and structured knowledge sharing.
  • Identify tooling gaps and build the business case for platform investments, working with engineering leadership to prioritize roadmap items.


Key Qualifications
  • 8+ years of DevOps or platform engineering experience, with at least 2 years operating at a Staff or Principal level in an organization of 100+ engineers.
  • Deep, hands-on expertise with Kubernetes - EKS specifically preferred - including troubleshooting workloads, networking, storage, and cluster operations at scale.
  • Strong command of Azure DevOps Pipelines, including YAML pipeline authoring, library management, service connections, and environment promotion gates.
  • Proven track record designing and maintaining CI/CD systems for microservice architectures with multiple independent teams as consumers.
  • Experience operating observability platforms (New Relic, Datadog, or similar) to drive proactive reliability improvements, not just reactive alerting.
  • Proficiency in at least one scripting language (Python, Bash, or Go) and Infrastructure-as-Code tooling (Terraform, Pulumi, or CDK).
  • Familiarity with feature flag patterns and operational considerations around progressive delivery (Unleash or equivalent is a plus).
  • Excellent written communication skills - you default to documentation and can translate complex infrastructure decisions into guidance engineers actually read.


Preferred Qualifications
  • Experience with data engineering or ML infrastructure workloads on Kubernetes (Spark on EKS, Argo Workflows, Airflow).
  • Background contributing to or maintaining internal developer portals (Backstage or similar).
  • Familiarity with FinOps practices and tooling for AWS cost attribution and optimization across shared Kubernetes clusters.
  • Experience in SRE-adjacent roles; comfort with SLO/SLI definition and error budget policy.


Special Work Conditions
  • Sedentary - Involves sitting most of the time but may involve walking or standing for brief periods of time. Some positions may entail exerting up to 15 lbs. of force occasionally and/or a negligible amount of force to lift, carry, push, or pull.


We take care of our people.
When you join Cast & Crew, you're backed by a benefits package built around what matters most. From comprehensive medical, dental, and vision coverage, to a 401(k) match, generous PTO, paid parental leave, health and wellness programs, tuition reimbursement, and employee discounts. We're committed to helping you thrive, professionally, and personally.

Benefits are subject to eligibility requirements.

Compensation is commensurate with various factors including, but not limited to, relevant experience, qualifications, skills, training, licensure, certifications, geographic cost of labor, and other business and organizational needs. Compensation range for candidates in other locations may differ based on the cost of labor in that location. The compensation range for this position is: $190,000.00 - $235,000.00 per year.

Similar Jobs

More Jobs at Cast & Crew

More Information Technology Jobs

Find similar Staff DevOps Engineer jobs: