Staff Infrastructure Engineer

Headway

$135K — $160K *
US-AnywhereRemote in San Francisco, CA
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in platform, infrastructure, or SRE roles with significant production traffic experience.
  • Expertise in AWS with hands-on production ownership of compute and networking.
  • Proficient in infrastructure-as-code, especially Terraform for self-serve platforms.
  • Experience with autoscaling and container orchestration using ECS and/or EKS.
  • Demonstrated ability to make deploys safe and self-serve for multiple teams.
  • Influence at Staff-level, capable of driving cross-team decisions effectively.

Responsibilities

  • Architect and manage the cloud platform used by all engineers at Headway.
  • Redesign deployment processes for improved failure containment and isolation.
  • Oversee ECS and EKS infrastructure and evaluate broader EKS integration for AI workloads.
  • Pre-compute capacity needs to manage fluctuating workloads more effectively.
  • Develop a self-serve infrastructure platform with necessary guardrails for engineering teams.
  • Establish detailed cost attribution mechanisms across AWS and other services.
  • Manage the health and performance of the Python runtime under load.

Benefits

  • Opportunities for professional growth in a fast-paced environment.
  • Work collaboratively with a technically advanced team.
  • Exposure to cutting-edge AI and healthcare technologies.
  • Contributions to a revolutionary mental healthcare system.
Full Job Description
The Role

Architect and own the cloud platform that every engineer at Headway deploys on. Make deploys boring, scaling automatic, infrastructure self-serve, and cost attributable.

About Infrastructure Engineering at Headway

Building a new mental healthcare system at Headway is only possible because of the scale and leverage that software can provide. Infrastructure Engineering is one of three teams in Engineering Foundations, alongside Agentic Engineering and Eddy, Headway's internal AI platform. We own the infrastructure Headway runs on, as well as the performance and scalability of Headway's FastAPI/Python monolith, and the AWS infrastructure it runs on. We are the paved road the rest of engineering builds on: databases, deploys, asynchronous platform, observability, autoscaling, and self-serve compute infrastructure.

As AI dramatically increases code velocity across the org, our mandate is to make every change safe from the start and build systems that get safer as they get faster.

About This Role

You will own the cloud platform Headway runs on. You will lead the work to isolate blast radius, make autoscaling trustworthy, and build a self-serve infrastructure platform so that the Infrastructure Engineering team focuses on the strategic and unusual, not the routine.

You will serve as the technical anchor for Headway's compute, networking, and deployment platform, and bring Staff-level influence to an area that every engineer depends on daily.

What You'll Own
  • Deployment architecture and blast-radius containment. Redesign deployment so that a mistake in one part of our service cannot block or take down others. Continue to drive our shift toward per-service deploy isolation and functional-area slices that contain failures rather than propagating them across the platform.
  • Container footprint and networking. Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design the next iteration of inter-service network connectivity.
  • Capacity and scaling. Own capacity for a spiky workload: floors computed ahead of demand rather than chased by reactive scaling, self-deriving from data, with drift caught early.
  • Self-serve infrastructure platform. Build the Terraform self-serve platform with guardrails so engineering teams own their standard infrastructure changes and Reliability Engineering reviews only the non-standard ones.
  • Cloud cost attribution and controls. Stand up per-team cost attribution across AWS, Datadog, and LLM spend. Make infrastructure costs visible and attributable so teams can make informed tradeoffs.
  • Python runtime and dependency health. Own how the monolith behaves under load: garbage collection, event loop contention, and the runtime limits that bite first. Lead the framework and package upgrades most teams defer.
Who You Are

A technical leader who combines deep cloud and infrastructure expertise with the instincts to drive alignment across teams. You thrive in ambiguity, operate with a high degree of ownership, and are as comfortable setting technical direction as you are building. You make other engineers better through architecture reviews, runbooks, and paved-road tooling that raises the baseline for everyone who ships on the platform.

Experience we're seeking:
  • 8 or more years in platform, infrastructure, or SRE roles at companies running significant production traffic
  • Deep AWS expertise and production ownership of compute and networking at scale (ECS, EKS, RDS, networking, IAM)
  • Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams
  • Hands-on autoscaling and capacity engineering, and container orchestration with ECS and/or EKS
  • Track record making deploys safe and self-serve for other teams, not just your own
  • Staff-level influence: you drive decisions across team boundaries and raise the infrastructure bar org-wide without requiring management authority to do it

Nice to have:
  • FinOps and cloud cost optimization experience
  • Kubernetes and EKS depth
  • Observability tooling at scale (Datadog)
  • Experience in healthcare or other regulated environments
  • Experience with event driven systems

Similar Jobs

More Jobs at Headway

More Information Technology Jobs

Find similar Staff Infrastructure Engineer jobs: