Software Engineer, Site Reliability

features and labels

$180K — $250K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years managing critical production systems and software development workflows
  • Strong production experience with Kubernetes at scale and infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking and container networking
  • Experience building CI/CD systems and GitOps workflows
  • Proficiency in Python and either Go or Bash for automation
  • Strong experience with logging, monitoring and alerting tools
  • Excellent communication skills and ability to drive technical decisions

Responsibilities

  • Own and operate Kubernetes infrastructure, including lifecycle, upgrades, and networking
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Leverage AI for automating production issue resolution and improving software reliability
  • Build dashboards, alerting, and anomaly detection systems
  • Define and enforce SLOs, and manage incident response processes
  • Improve networking, load balancing, and service mesh configurations
  • Drive reliability improvements through automation and chaos engineering

Benefits

  • Challenging and interesting work
  • Learning and growth opportunities
  • Health, dental, and vision insurance
  • Regular team events and offsites
  • Located in downtown San Francisco
Full Job Description
About this role

You are a seasoned SRE who keeps production infrastructure running at scale. You own the reliability and availability of customer-facing systems - from Kubernetes clusters to deployment pipelines to the networking layer that connects it all. You think in SLOs, automate ruthlessly, and treat every incident as a chance to make the system better.
Key Responsibilities
  • Own and operate our Kubernetes infrastructure: cluster lifecycle, upgrades, networking, and multi-tenant isolation for customer workloads
  • Build and maintain CI/CD pipelines and deployment infrastructure
  • Leverage AI to an extreme level to automate analysis and resolution of production issues, and improve software development speed, reliability and maintainability
  • Build dashboards, alerting, and anomaly detection across our systems
  • Define and enforce SLOs and build out incident response processes
  • Manage and improve our networking, load balancing, and service mesh configurations
  • Drive reliability improvements across the stack through automation, runbooks, and chaos engineering
Requirements
  • 5+ years experience in managing critical production systems and software development workflows
  • Strong production experience setting up and operating Kubernetes at scale, using infrastructure-as-code (Terraform, Ansible)
  • Deep knowledge of Linux networking, container networking (CNI plugins, VXLAN, BGP), and DNS
  • Experience building CI/CD systems and GitOps workflows (FluxCD, ArgoCD)
  • Proficiency in Python and either Go or Bash for tooling and automation
  • Strong experience with logging, monitoring and alerting (Prometheus, Grafana, Loki, Thanos, VictoriaMetrics, Datadog)
  • Excellent communication and ability to drive technical decisions across teams
  • Self-starter who executes quickly, takes ownership, and constantly seeks improvement
Nice to have
  • Experience with managing GPU and AI/ML workloads
  • Experience with kernel-based monitoring and routing (eBPF, XDP)
  • Experience with security tooling (Falco, Coroot, SIEM)
  • Experience with bare metal Kubernetes networking (Calico, Cilium, MetalLB)
  • Experience with distributed storage systems (Ceph, Longhorn, etc.)
Compensation
  • $180,000-250,000 plus equity + benefits (Range is based across 3 levels MId, Senior and Staff)
Location
  • San Francisco, CA

What we offer at fal
  • Interesting and challenging work
  • A lot of learning and growth opportunities
  • We are currently hiring in downtown San Francisco.
  • Health, dental, and vision insurance (US)
  • Regular team events and offsites

Similar Jobs

More Jobs at features and labels

More Information Technology Jobs

Find similar Software Engineer, Site Reliability jobs: