Site Reliability Engineer

Astera

• $130K — $155K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in site reliability engineering or a related field
  • Strong understanding of systems architecture including schedulers, containers, networking, storage, and hardware
  • Ability to manage high-pressure situations and take ownership of system reliability
  • Valuable experience with observability tools and operational standards
  • Familiarity with modern tech stacks, including Docker, Kubernetes, and Python, though flexibility is encouraged

Responsibilities

  • Ensure efficient access to compute resources for researchers
  • Provide transparency into resource utilization and health of clusters
  • Facilitate automatic scaling of resources in response to demand
  • Manage access permissions to guarantee appropriate resource allocation
  • Drive reproducibility in research environments through deterministic deployments
  • Automate processes to enhance operational efficiency

Benefits

  • Opportunities to work at the intersection of neuroscience and AI
  • Autonomy in adapting infrastructure to meet diverse research needs
  • Access to cutting-edge computational resources
  • Collaborative culture with a focus on creativity and innovation
  • Potential for professional growth in a startup-like environment
Full Job Description
Position Summary

We are looking for a Site Reliability Engineer to own the digital infrastructure that powers our research.

This includes compute resources that we rent from third parties, container registries, and dashboards. The main objective is to make sharing these resources easy and efficient, ensuring the infrastructure is reliable and accessible to the right people.

This role spans a broad spectrum of activities:
  • Compute Access: Ensure easy and efficient access to compute resources for our researchers.
  • Resource Visibility: Provide clear visibility into resource utilization and cluster health.
  • Auto-Scaling: Enable automatic scaling of compute resources based on demand.
  • Access Management: Ensure the right people have access to the right resources.
  • Reproducibility: Drive towards deterministic deployments and reproducible research environments.
  • Process Automation: Automate operational processes where it makes sense to increase efficiency.
  • Current stack: Ansible, Kubernetes, Docker, Tailscale, Python, Grafana, Prometheus, and Talos Linux. We're not religious about any of it.


Qualifications
  • Ownership: You are comfortable being the person accountable when the cluster is unhealthy or capacity is tight.
  • Systems Intuition: You understand how schedulers, containers, networking, storage, and hardware interact. You can reason about failure modes and design systems that degrade predictably.
  • Operational Rigor: You value observability, reproducibility, and clear operational boundaries. You leave systems in a state that other engineers can understand, operate, and debug without you.
  • Pragmatism: You can support experimental research workloads without forcing everything into a rigid "production" mold. You know when to stabilize and when to allow controlled chaos to speed up discovery.


Location & Visa
  • This role is in-person in Emeryville, CA.
  • Visa sponsorship may be available for qualified candidates.

Similar Jobs

More Jobs at Astera

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: