Staff Site Reliability Engineer

SoundHound AI Inc

$120K — $150K *
Information Technology
11 - 15 years of experience
Job Overview by Ladders

Qualifications

  • 12+ years of software engineering experience, particularly in Site Reliability Engineering or DevOps roles.
  • Expert-level proficiency in Google Cloud Platform (GCP) services such as GKE and Cloud Run.
  • Skilled in using Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Extensive knowledge of Kubernetes and service mesh architectures.
  • Strong experience with monitoring and observability tools like Datadog and Prometheus.
  • Proven track record in designing and managing high-throughput distributed systems.
  • Exceptional problem-solving abilities and effective communication skills, especially in mentoring.

Responsibilities

  • Design and maintain scalable infrastructure on Google Cloud Platform.
  • Automate CI/CD pipelines for reliable deployments.
  • Implement monitoring and observability strategies for system issue resolution.
  • Collaborate with engineering teams to optimize backend services.
  • Drive incident response and long-term remediation efforts.
  • Identify and eliminate operational inefficiencies.
  • Lead compliance initiatives within the department.

Benefits

  • Equity and comprehensive healthcare options.
  • Paid time off to promote work-life balance.
  • Remote work flexibility throughout Canada.
Full Job Description
The Opportunity

We're looking for a Staff Software Engineer (SRE) to join our Retail and Restaurants AI team. You will be responsible for the reliability, scalability, and performance of our infrastructure, with a deep focus on Google Cloud Platform (GCP). You will architect and maintain high-availability systems, automate operational tasks, and ensure our services can handle the demands of millions of voice AI interactions.

What You'll Do

  • Design, build, and maintain highly available and scalable infrastructure on Google Cloud Platform.
  • Architect and automate CI/CD pipelines to ensure rapid, reliable deployments.
  • Implement robust monitoring, alerting, and observability strategies to proactively identify and resolve system issues.
  • Partner with engineering teams to optimize performance, cost, and reliability of backend services.
  • Drive incident response, post-mortem analysis, and long-term remediation efforts.
  • Identify and eliminate sources of toil, promoting operational maturity and self-service capabilities.
  • Collaborate with cross-functional teams to ensure alignment on infrastructure roadmaps and security standards.
  • Lead department wide compliance (PCI, SOC) initiatives.


What You'll Bring

  • 12+ years of software engineering experience, with significant experience in Site Reliability Engineering or DevOps roles.
  • Expert-level experience with Google Cloud Platform (GCP) services (e.g., GKE, Compute Engine, Cloud Run, Pub/Sub).
  • Proficient in Infrastructure as Code (IaC) tools like Terraform or Pulumi.
  • Deep experience with Kubernetes, container orchestration, and service mesh architectures.
  • Strong background in monitoring and observability tools (e.g., Datadog, Prometheus, Grafana, Cloud Monitoring).
  • Experience designing and managing high-throughput, distributed systems.
  • Strong problem-solving skills and a growth mindset-comfortable with ambiguity and making high-stakes technical trade-offs.
  • Excellent communication skills and a demonstrated ability to mentor engineers.


Preferred Qualifications

  • Experience working in a high-velocity, customer-focused environment.
  • Familiarity with functional programming paradigms (e.g., Clojure/ClojureScript).
  • Prior experience in the restaurant technology, hospitality, or AI-driven SaaS space.
  • Experience implementing security and compliance best practices in the cloud.


Workplace & Compensation

This role is available throughout Canada.

Compensation includes salary, equity, comprehensive healthcare, paid time off, and other benefits. Our recruiting team will provide a specific salary range based on location and years of experience.

#LI-MQ1 #LI-REMOTE

Let's Start the Conversation

Join SoundHound AI and collaborate with colleagues worldwide who are shaping the future of voice AI. Guided by our values-supportive, open, undaunted, nimble, and determined to win-we strive to build breakthrough AI experiences together.

Similar Jobs

More Jobs at SoundHound AI Inc

  • Gem
    Global Equity Manager
    $100K — $150K *
    Remote
    Finance & Insurance
    Remote in United States
  • Gem
    Global Equity Manager
    $100K — $150K *
    Santa Clara, CA 95051 (Santa Clara County)
    Finance & Insurance
    In-Person

More Information Technology Jobs

Find similar Staff Site Reliability Engineer jobs: