The RoleArchitect and own the cloud platform that every engineer at Headway deploys on. Make deploys boring, scaling automatic, infrastructure self-serve, and cost attributable.
About Infrastructure Engineering at HeadwayBuilding a new mental healthcare system at Headway is only possible because of the scale and leverage that software can provide. Infrastructure Engineering is one of three teams in Engineering Foundations, alongside Agentic Engineering and Eddy, Headway's internal AI platform. We own the infrastructure Headway runs on, as well as the performance and scalability of Headway's FastAPI/Python monolith, and the AWS infrastructure it runs on. We are the paved road the rest of engineering builds on: databases, deploys, asynchronous platform, observability, autoscaling, and self-serve compute infrastructure.
As AI dramatically increases code velocity across the org, our mandate is to make every change safe from the start and build systems that get safer as they get faster.
About This RoleYou will own the cloud platform Headway runs on. You will lead the work to isolate blast radius, make autoscaling trustworthy, and build a self-serve infrastructure platform so that the Infrastructure Engineering team focuses on the strategic and unusual, not the routine.
You will serve as the technical anchor for Headway's compute, networking, and deployment platform, and bring Staff-level influence to an area that every engineer depends on daily.
What You'll Own- Deployment architecture and blast-radius containment. Redesign deployment so that a mistake in one part of our service cannot block or take down others. Continue to drive our shift toward per-service deploy isolation and functional-area slices that contain failures rather than propagating them across the platform.
- Container footprint and networking. Own the ECS and EKS footprint, evaluate broader EKS adoption for AI workloads, and design the next iteration of inter-service network connectivity.
- Capacity and scaling. Own capacity for a spiky workload: floors computed ahead of demand rather than chased by reactive scaling, self-deriving from data, with drift caught early.
- Self-serve infrastructure platform. Build the Terraform self-serve platform with guardrails so engineering teams own their standard infrastructure changes and Reliability Engineering reviews only the non-standard ones.
- Cloud cost attribution and controls. Stand up per-team cost attribution across AWS, Datadog, and LLM spend. Make infrastructure costs visible and attributable so teams can make informed tradeoffs.
- Python runtime and dependency health. Own how the monolith behaves under load: garbage collection, event loop contention, and the runtime limits that bite first. Lead the framework and package upgrades most teams defer.
Who You AreA technical leader who combines deep cloud and infrastructure expertise with the instincts to drive alignment across teams. You thrive in ambiguity, operate with a high degree of ownership, and are as comfortable setting technical direction as you are building. You make other engineers better through architecture reviews, runbooks, and paved-road tooling that raises the baseline for everyone who ships on the platform.
Experience we're seeking:
- 8 or more years in platform, infrastructure, or SRE roles at companies running significant production traffic
- Deep AWS expertise and production ownership of compute and networking at scale (ECS, EKS, RDS, networking, IAM)
- Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams
- Hands-on autoscaling and capacity engineering, and container orchestration with ECS and/or EKS
- Track record making deploys safe and self-serve for other teams, not just your own
- Staff-level influence: you drive decisions across team boundaries and raise the infrastructure bar org-wide without requiring management authority to do it
Nice to have:
- FinOps and cloud cost optimization experience
- Kubernetes and EKS depth
- Observability tooling at scale (Datadog)
- Experience in healthcare or other regulated environments
- Experience with event driven systems