Staff Infrastructure Engineer

Headway

$138K — $165K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in platform, infrastructure, or SRE roles with significant production traffic experience.
  • Deep AWS expertise including ECS, EKS, RDS, networking, IAM.
  • Strong infrastructure-as-code experience, particularly with Terraform.
  • Hands-on experience with autoscaling and container orchestration.
  • Proven ability to make deploys safe and self-serve for multiple teams.
  • Demonstrated staff-level influence across team boundaries.

Responsibilities

  • Architect and own the cloud platform supporting all engineering deployments.
  • Redesign deployment architecture to isolate service failures.
  • Manage ECS and EKS footprint and enhance network connectivity for AI workloads.
  • Develop capacity planning for variable workloads to preemptively address demand.
  • Create a self-serve infrastructure platform with defined guardrails.
  • Implement cost attribution across cloud services for informed decision-making.
  • Oversee the performance of the Python runtime under load and manage upgrades.

Benefits

  • Flexible work environment with a focus on work-life balance.
  • Opportunities for professional development and growth.
  • Collaborative and inclusive team culture.
  • Access to cutting-edge technology and tools.
  • Impactful work in the mental healthcare space.
Full Job Description
The Role

Architect and own the cloud platform that every engineer at Headway deploys on. Make deploys boring, scaling automatic, infrastructure self-serve, and cost attributable.

About Infrastructure Engineering at Headway

Building a new mental healthcare system at Headway is only possible because of the scale and leverage that software can provide. Infrastructure Engineering is one of three teams in Engineering Foundations, alongside Agentic Engineering and Eddy, Headway's internal AI platform. We own the infrastructure Headway runs on, as well as the performance and scalability of Headway's FastAPI/Python monolith, and the AWS infrastructure it runs on. We are the paved road the rest of engineering builds on: databases, deploys, asynchronous platform, observability, autoscaling, and self-serve compute infrastructure.

As AI dramatically increases code velocity across the org, our mandate is to make every change safe from the start and build systems that get safer as they get faster.

About This Role

You will own the cloud platform Headway runs on. You will lead the work to isolate blast radius, make autoscaling trustworthy, and build a self-serve infrastructure platform so that the Infrastructure Engineering team focuses on the strategic and unusual, not the routine.

You will serve as the technical anchor for Headway's compute, networking, and deployment platform, and bring Staff-level influence to an area that every engineer depends on daily.

What You'll Own
  • Deployment architecture and blast-radius containment. Redesign deployment so that a mistake in one part of our service cannot block or take down others. Continue to drive our shift toward per-service deploy isolation and functional-area slices that contain failures rather than propagating them across the platform.
  • Container footprint and networking. Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design the next iteration of inter-service network connectivity.
  • Capacity and scaling. Own capacity for a spiky workload: floors computed ahead of demand rather than chased by reactive scaling, self-deriving from data, with drift caught early.
  • Self-serve infrastructure platform. Build the Terraform self-serve platform with guardrails so engineering teams own their standard infrastructure changes and Reliability Engineering reviews only the non-standard ones.
  • Cloud cost attribution and controls. Stand up per-team cost attribution across AWS, Datadog, and LLM spend. Make infrastructure costs visible and attributable so teams can make informed tradeoffs.
  • Python runtime and dependency health. Own how the monolith behaves under load: garbage collection, event loop contention, and the runtime limits that bite first. Lead the framework and package upgrades most teams defer.
Who You Are

A technical leader who combines deep cloud and infrastructure expertise with the instincts to drive alignment across teams. You thrive in ambiguity, operate with a high degree of ownership, and are as comfortable setting technical direction as you are building. You make other engineers better through architecture reviews, runbooks, and paved-road tooling that raises the baseline for everyone who ships on the platform.

Experience we're seeking:
  • 8 or more years in platform, infrastructure, or SRE roles at companies running significant production traffic
  • Deep AWS expertise and production ownership of compute and networking at scale (ECS, EKS, RDS, networking, IAM)
  • Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams
  • Hands-on autoscaling and capacity engineering, and container orchestration with ECS and/or EKS
  • Track record making deploys safe and self-serve for other teams, not just your own
  • Staff-level influence: you drive decisions across team boundaries and raise the infrastructure bar org-wide without requiring management authority to do it

Nice to have:
  • FinOps and cloud cost optimization experience
  • Kubernetes and EKS depth
  • Observability tooling at scale (Datadog)
  • Experience in healthcare or other regulated environments
  • Experience with event driven systems

Similar Jobs

More Jobs at Headway

More Information Technology Jobs

Find similar Staff Infrastructure Engineer jobs: