Site Reliability Engineer

Picogrid

$170K — $195K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years in site reliability, or similar roles
  • Deep knowledge of Kubernetes operations
  • Experience in building observability dashboards
  • Proficient in incident response and postmortems
  • Hands-on with Terraform or OpenTofu
  • Fluent in AWS management and hardening
  • Experience in high-availability database operations
  • Familiarity with IoT or edge fleet management
  • Comfortable in fast-paced, uncertain environments

Responsibilities

  • Drive and shape reliability SLIs and SLOs for cloud systems
  • Establish reliability SLIs and SLOs for edge devices in challenging environments
  • Manage the observability stack including Grafana and Prometheus
  • Engage in on-call duties and conduct incident response
  • Implement infrastructure as code to enhance reliability

Benefits

  • Significant stock options offered
  • 401(k) with employer matching available
  • Comprehensive health coverage including dental and vision
  • Relocation assistance for eligible candidates
  • Unlimited PTO with a minimum of two weeks off
  • Paid parental leave for all parents
  • In-office lunches and a stocked kitchenette
  • EV charging provided at HQ
  • Unique and stimulating office environment
Full Job Description

About the Role

As Picogrid's first Site Reliability Engineer you will own production reliability across cloud and edge, from observability and incident response through node lifecycle, stateful workloads, and a fleet of hardware edge devices in the field. You will help build and define the systems, processes and best practices that ensure Picogrid's systems can be relied upon by our warfighters in even the toughest battlefield conditions. You will work with engineers to build a strong on-call culture where issues are root caused swiftly, and ensure our alerting and monitoring have exceptional coverage and signal-to-noise ratio.

Security is a shared responsibility across all our DevSecOps roles, and as part of a scrappy startup team you will be expected to help stand up new infrastructure and other related DevSecOps tasks as needed.

Responsibilities
  • Own, define and drive our reliability SLIs and SLOs for cloud deployments
  • Own, define and drive our reliability SLIs and SLOs for our edge devices deployed in remote and sometimes contested areas
  • Own the observability stack: Grafana, Prometheus, Loki, and OpenTelemetry, with dashboards versioned in git and alerting rules checked in alongside the code they watch
  • Participate in on-call and incident response: log-first troubleshooting, blameless postmortems, and follow-up hardening
  • Encode reliability into infrastructure as code


Required Qualifications
  • 3+ years of experience as an SRE or related roles
  • Deep Kubernetes operations experience: node lifecycle, workload scheduling, StatefulSets, graceful drains, and live cluster debugging
  • Experience designing comprehensive observability dashboards and high signal-to-noise ratio alerting rules
  • You are a competent and experienced incident responder practicing methodical evidence-first triage, blameless postmortems, and turning incidents into durable guardrails
  • Production Terraform or OpenTofu experience
  • Fluent in AWS including IAM, networking, multi-account environments, and account and workload hardening
  • Experience managing high availability database deployments
  • IoT or edge fleet operation experience
  • Comfortable operating in scrappy, fast-paced environments, and turning ambiguous requirements into concrete solutions
  • You optimize for providing value early in projects and short iteration cycles


Preferred Qualifications
  • GovCloud, FIPS, or other regulated or air-gapped environment experience
  • Constrained edge hardware such as NVIDIA Jetson platforms (AGX Thor, Orin Nano), including shared CPU and GPU memory and thermal constraints
  • Overlay or mesh networking operations: Nebula, WireGuard, Tailscale, or similar
  • Standing up SLO and error-budget tooling (sloth, Pyrra, or equivalent) from scratch
  • Active security clearance


Compensation & Benefits
  • Base salary range: $170,000 - $195,000 per year. Base salary is just one part of your total compensation package at Picogrid.
  • Significant stock options with a high potential upside as an early-stage company
  • 401(k) with employer matching
  • Full health coverage (medical, dental, and vision insurance)
  • Relocation assistance provided (if applicable)
  • Unlimited PTO (two-week minimum) and 11 paid holidays per year
  • Paid parental leave for both parents
  • Lunch provided when working in-office and a fully stocked kitchenette
  • Free EV charging at the HQ
  • Unique office in El Segundo, CA stocked with quality coffee, snacks, and craft beer


Export Control Requirements

To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. a7 1157, or (iv) Asylee under 8 U.S.C. a7 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.

Similar Jobs

More Jobs at Picogrid

  • Strategy & Operations Manager
    $130K — $190K *
    El Segundo, CA 90245 (Los Angeles County)
    Aerospace & Defense
    In-Person
  • Developer Experience Engineer
    $170K — $195K *
    El Segundo, CA 90245 (Los Angeles County)
    Information Technology
    In-Person
  • Site Reliability Engineer
    $170K — $195K *
    El Segundo, CA 90245 (Los Angeles County)
    Information Technology
    In-Person
  • Head of People
    $160K — $200K *
    El Segundo, CA 90245 (Los Angeles County)
    Business Services
    In-Person
  • Executive Assistant
    $90K — $130K *
    El Segundo, CA 90245 (Los Angeles County)
    Aerospace & Defense
    In-Person

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: