5-7 years in software engineering or site reliability engineering roles
Proficient at building reliability tooling and automation from scratch
Deep understanding of observability and incident response practices
Experience with Kubernetes and cloud infrastructure
Ability to diagnose issues across application and infrastructure layers
Familiarity with distributed systems and real-time networking is advantageous
Responsibilities
Learn Vapi's architecture and operational reliability within the first 30 days
Take ownership of a reliability workstream focusing on observability and automation by 60 days
Contribute to initial reliability improvements quickly after joining
Drive down manual processes and enhance team responsiveness to system failures
Become the key holder of a significant part of reliability and propose future enhancements by 90 days
Benefits
Comprehensive health coverage including medical, dental, and vision
Quarterly off-sites for team bonding
Flexible time off policy allowing employees to take what they need
Enjoy catered meals and transportation perks
Annual $10k budget for learning and development activities
Full Job Description
Why this Role:
Vapi's real-time voice platform depends on core infrastructure across compute, storage, networking, and telephony. As usage grows, reliability has to be designed into the systems our customers depend on.
We're adding a dedicated SRE profile to our five-person Infrastructure team. We need a software-first engineer who can build tooling and automation, improve observability, and turn operational lessons into lasting engineering improvements.
This role has a distinct focus from our broader Infrastructure Engineer role because it requires deep reliability and operability judgment. You'll join a calibrated team and work on latency-sensitive systems at meaningful scale. This role is based in San Francisco.
What You'll Do:
30 Days: Learn Vapi's architecture, production environment, incident history, and reliability practices. Build context with the Infrastructure and product engineering teams, then contribute an initial operational or reliability improvement.
60 Days: Own a reliability workstream across observability, incident response, capacity, performance, or production automation. Reduce manual work and improve how the team detects, understands, and responds to failures.
90 Days: Become the go-to owner for a meaningful part of Vapi's reliability surface. Deliver a durable improvement to failure prevention or recovery, and propose a roadmap for the next reliability investments.
Who You Are:
You are a senior or staff-level software engineer with meaningful SRE, production engineering, or infrastructure experience in distributed systems.
You write production-quality software and have built reliability tooling or automation yourself, rather than relying only on operational process.
You have deep experience with observability, incident response, failure analysis, capacity, and the practices that keep production systems healthy.
You are comfortable with Kubernetes, networking, and cloud infrastructure, and you can debug across application and infrastructure boundaries.
You reason clearly about failure modes and can balance reliability investments with product and engineering velocity.
Experience with real-time networking or telephony, Envoy, Postgres, Redis, Kafka, Aurora, ClickHouse, or a Google-style SRE environment is a strong plus.
How We Work:
Build something worthy of love
Craft matters. We aim to build products and experiences customers genuinely love, not just tolerate.
Commit and follow through
We finish what we start and build trust by being people others can count on.
Why not today?
We value urgency and momentum. The fastest path to customer value usually wins.
Seek raw input
We go directly to customers, data, and teammates instead of relying on summaries or assumptions.
It's our problem
We operate as one team. We share credit, own mistakes together, and support each other when things get hard.
Be direct and kind
We give feedback clearly, respectfully, and without delay.
Why Vapi:
Generational impact: Build the human interface for every business
Ownership culture: 70% of the company are previous founders
Kind team: The founders, Jordan and Nikhil, are Canadians
Tier-1 Investors: YC, KP seed, Bessemer Series A
What We Offer:
Real stake: We offer a base salary of $280,000 to $314,000 and excellent equity ownership
Comprehensive health coverage: medical, dental, and vision plans
Team love: We love hanging out, and we do quarterly off-sites
Flexible time off: take what you need
More: catered meals, transportation, gym, and a $10k annual L&D budget