5-7 years of experience in CI/CD or release infrastructure at scale.
Strong knowledge of distributed systems and diagnosing failures across multiple layers.
Hands-on experience with Kubernetes-based systems and tools like Argo CD, Terraform, and Atlantis.
Proficient in automating workflows and enhancing system reliability with data-driven approaches.
Curious with the ability to adapt to changing architectures and products swiftly.
Effective communicator with a team-oriented mindset; SRE experience and familiarity with AI coding tools are advantages.
Responsibilities
Map and assess Vapi’s release architecture in partnership with product and infrastructure teams.
Manage the release-tooling roadmap and enhance its operational resilience.
Eliminate costly bottlenecks in CI/deployment processes and increase observability.
Create a repeatable model for safe and efficient releases with measurable improvements.
Implement self-service workflows and set long-term release infrastructure goals.
Benefits
Comprehensive health coverage including medical, dental, and vision plans.
Flexible time off policy to accommodate personal needs.
Regular team bonding activities including quarterly off-sites.
Additional perks like catered meals, transportation support, and a $10,000 annual learning and development budget.
Full Job Description
Why Were Hiring This Role:
Vapi ships hundreds of pull requests a day while serving some of the worlds largest enterprises. Every release must stay fast, safe, observable, and predictable at scale.
Our CI and deployment systems are critical product infrastructure. We need a senior/staff engineer to remove bottlenecks, repair fragile deploy paths, and increase confidence without slowing product teams.
Youll own the systems behind safe releases, including Argo CD, Terraform, and Atlantis-and automate the reliable path. This role focuses on CI/CD, testing, and deployment safety, not primary incident response.
What Youll Do:
30 Days: Map Vapis release architecture end to end. Partner with product and infrastructure engineers, baseline CI/test/deploy performance, identify the biggest failure modes, and ship an initial improvement.
60 Days: Own the release-tooling roadmap and operating health. Remove the costliest CI/deploy bottlenecks, improve observability and diagnosis, and automate slow or failure-prone workflows.
90 Days: Establish a repeatable model for safe, fast releases with measurable gains in validation time, deploy speed, and reliability. Deliver self-service workflows and guardrails, and set the longer-term release-infrastructure roadmap.
Who You Are:
Evidence of owning critical CI/CD or release infrastructure at meaningful scale and driving measurable cross-team improvements. We care about impact, not a specific number of years.
Strong distributed-systems fundamentals and the ability to diagnose failures across build, test, deployment, and infrastructure layers.
Hands-on experience operating Kubernetes-based delivery systems and infrastructure-as-code tooling such as Argo CD, Terraform, and Atlantis or comparable systems.
A habit of automating repetitive work, investing in testing and passive defenses, and using data to prove that a system became faster or more reliable.
Deep curiosity and the ability to gather context quickly, form a clear point of view, and keep pace as the architecture and product evolve.
A low-ego teammate who communicates clearly and owns outcomes. Familiarity with AI coding tools is expected; experience as an SRE, with development frameworks like multi-agent workflows, and programming languages like TypeScript and Go are a plus.
This is a full-time, hybrid role in San Francisco.
How We Work:
Build something worthy of love
Craft matters. We aim to build products and experiences customers genuinely love, not just tolerate.
Commit and follow through
We finish what we start and build trust by being people others can count on.
Why not today?
We value urgency and momentum. The fastest path to customer value usually wins.
Seek raw input
We go directly to customers, data, and teammates instead of relying on summaries or assumptions.
Its our problem
We operate as one team. We share credit, own mistakes together, and support each other when things get hard.
Be direct and kind
We give feedback clearly, respectfully, and without delay.
Why Vapi:
Generational impact: Build the human interface for every business
Ownership culture: Many of us are previous founders
Kind team: The founders, Jordan and Nikhil, are Canadians
Tier-1 Investors: YC, KP seed, Bessemer Series A
What We Offer:
Real stake: $235,000-$280,000 base salary, plus meaningful equity ownership
Comprehensive health coverage: medical, dental, and vision plans
Team love: We love hanging out, and we do quarterly off-sites
Flexible time off: take what you need
More: catered meals, transportation, equipment, and a $10k annual L&D budget