Senior Software Engineer, Reliability

Tin Can

$130K — $155K *
Telecommunications & Hardware
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience with production backend or distributed systems
  • Strong proficiency with AWS and Infrastructure as Code (Terraform)
  • Experience owning reliability practices, including SLOs and alerting
  • Familiarity with telephony systems and protocols like SIP and RTP
  • Ability to utilize AI in diagnosis and analysis
  • Hands-on problem-solving approach with willingness to dive deep into issues
  • Effective communicator with strong documentation skills
  • Collaborative mindset, willing to contribute across the team

Responsibilities

  • Own the reliability outcomes for family phone calls and activation flows
  • Implement change management for customer-visible features
  • Manage production pipelines for application and call-path changes
  • Ensure staging tier tests against production-like traffic
  • Conduct load tests for peak holiday usage
  • Build observability around product outcomes and infrastructure metrics
  • Drive root cause analysis across the technology stack
  • Oversee incident reviews and ensure accountability for failures
  • Collaborate with PM on engineering backlog based on diagnostics

Benefits

  • Collaborative work environment that values contributions from all team members
  • Opportunities to work with cutting-edge technology in telephony and infrastructure
  • Focus on making meaningful impacts in family experiences with technology
  • Emphasis on professional development and growth within the engineering discipline
  • Chance to innovate in a dynamic, consumer-focused setting
Full Job Description
The Role

A Tin Can has exactly one job. A kid picks it up, and the person they want to talk to is on the other end - clear, and right away, every time. When that doesn't happen we haven't shipped a bug, we've broken the promise a family bought us for. There is no feature we could ship this year that matters more than getting this right.

We're looking for a Senior Software Engineer to own that outcome end to end. Our calls run on FreeSWITCH and Kamailio in AWS, gated by a set of Lambda services, with an activation flow every family walks through the morning they open the box. You'll own how that path is built, tested, deployed, and measured: defined as code, proven in staging before it reaches a kid's phone, and instrumented so we see a problem the moment it starts. You'll work alongside a VoIP specialist who owns the SIP and media plane itself, an SRE who owns the platform underneath, our firmware engineers, and the PM who owns reliability. Your job is the outcome across all of it, and you'll go as deep as any layer requires.
What You'll Do
  • Own the reliability outcome families actually experience - did the call connect, did it sound right, was the phone reachable - and the number we hold ourselves to
  • Make change management real for customer-visible behavior, so a migration that changes what a family hears is announced rather than discovered
  • Own the pipelines that carry an application or call-path change into production: reviewed, tested, and reversible
  • Own the staging tier as a real gate, where a change proves itself against production-like traffic before a kid's phone sees it
  • Load-test the activation chain ahead of our holiday peak, so the busiest morning of the year is one we've already rehearsed
  • Build the product-outcome layer of our observability - the funnels and SLOs showing whether a call, voicemail, or activation actually succeeded - alongside the infrastructure metrics our SRE owns
  • Drive incidents to real root cause across the stack, from SIP signaling and media to Lambda services and Postgres
  • Own the incident-review practice: every class of failure gets an owner, a runbook, and a test that would catch it next time
  • Partner with our reliability PM to turn diagnosis into a backlog engineering can actually execute against
What We're Looking For
  • 5+ years building and operating production backend or distributed systems where downtime is something real people notice
  • Real depth in AWS and infrastructure as code (we use Terraform), and a bias toward reproducibility
  • You've owned reliability as a discipline, not a fire drill - SLOs, meaningful alerting, load testing, and incident review that changes what gets built next
  • Comfortable in real-time and telephony systems, or genuinely eager to get there; SIP, RTP, and media behavior are the substrate here
  • Fluency using AI as a multiplier across diagnosis, analysis, and testing
  • Hands-on and scrappy; you'd rather run the load test or read the SIP trace yourself than wait on someone else
  • A clear communicator who writes things down
  • A collaborative, low-ego approach - everyone here sweeps the floor
  • Bonus: production FreeSWITCH, Kamailio, or carrier SIP trunking; load-testing a system through a seasonal peak; Postgres at scale; operating a consumer hardware fleet; or observability tooling like Grafana

If you're excited by the idea of making a kid's phone call something their family never has to think about, we'd love to meet you.

Similar Jobs

More Jobs at Tin Can

More Telecommunications & Hardware Jobs

Find similar Senior Software Engineer, Reliability jobs: