Ooma

Site Reliability Engineer

Ooma$120K — $170K *
US-AnywhereRemote in United States
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 8+ years in Site Reliability Engineering or related fields, focusing on production systems and service delivery.
  • Extensive experience with on-premises data centers and managing bare-metal servers and VMs.
  • Deep knowledge of Linux operating systems, including troubleshooting and performance tuning.
  • Hands-on experience with server hardware lifecycle management and out-of-band management.
  • Strong understanding of Linux networking, including VLANs and DNS management.
  • Proven expertise in containers and orchestration, particularly Kubernetes.
  • Experience in managing CI/CD pipelines using tools like GitOps and Argo CD.

Responsibilities

  • Guide management of large data centers with hundreds of bare-metal servers and VMs.
  • Monitor and troubleshoot performance and reliability using observability tools.
  • Manage the server hardware lifecycle including procurement and capacity planning.
  • Administer virtualization and storage across global data centers.
  • Automate provisioning and OS lifecycle for efficiency.
  • Work with data center network infrastructure, handling latency and throughput issues.
  • Design and maintain scalable infrastructure using Kubernetes and microservices.

Benefits

  • Comprehensive Medical/Dental/Vision insurance for employees and dependents.
  • Employer Paid Income Protection Benefits including Life and Disability.
  • Flexible Spending Accounts for healthcare and dependent care.
  • 401(k) plan with employer matching contributions.
  • Employee Stock Purchase Plan for stock acquisition.
  • Paid Time Off along with sick leave and designated holidays.
  • Employee Assistance Program to support personal well-being.
Full Job Description
About the Role:

As a Site Reliability Engineer, you will leverage your extensive expertise in Linux systems, virtualization, containers, Kubernetes clusters, and CI/CD pipelines to ensure the stability and efficiency of our systems, collaborating across teams to implement best practices for infrastructure management, automated deployment, and application performance monitoring.

Deep on-premises experience is a core requirement, not a secondary consideration. Our production environment runs on our own hardware - large data centers built on hundreds of bare metal servers and VMs, with our own storage, virtualization, and physical network beneath them. You will operate comfortably at the hardware and OS layer while also owning the container and delivery platform on top of it.

**Location and Onsite Requirement: This role requires onsite work at least once per week at one of our designated data centers in Dallas, TX, Ashburn, VA or San Jose, CA. Candidates must be able to commute regularly to one of these locations. Relocation assistance and reimbursement for routine commuting, travel, or overnight lodging are not available.

What You'll Do:
  • Provide expert guidance on managing large data centers, including hundreds of bare metal servers and virtual machines (VMs), ensuring optimal configuration and performance.
  • Monitor and troubleshoot system performance, reliability, and availability using modern observability tools and techniques, with strong emphasis on diagnosing and resolving issues in operating systems and bare metal environments.
  • Manage the server hardware lifecycle - firmware and BIOS baselines, out-of-band management, failure triage, vendor and RMA coordination, capacity forecasting, and refresh planning.
  • Administer virtualization platforms and manage storage appliances across all data center locations worldwide.
  • Automate bare metal and VM provisioning and OS lifecycle at scale.
  • Work hands-on with data center network infrastructure - VLANs, routing, link aggregation, load balancers, and firewall rules - troubleshooting latency, throughput, and packet loss from the Linux host outward, including multi-site and colocation resiliency.
  • Design, implement, and maintain scalable, reliable infrastructure using containers, Kubernetes, and microservices architecture, including clusters on bare metal and on-premises VMs.
  • Oversee configuration management for consistent, reliable releases across environments, using Ansible for system configuration, patch management, and provisioning across data center infrastructure, and eliminating configuration drift across the fleet.
  • Design and operate high-throughput Kafka clusters for event streaming - topics, partitions, replication, consumer lag monitoring, and disaster recovery across data center infrastructure.
  • Implement name services and server management practices, including DNS, DHCP, NTP, directory and authentication services, and certificate management.
  • Collaborate with development teams to influence system design choices and operational policies, and continuously evaluate and integrate new technologies, hardware platforms, and automation approaches to improve operational efficiency and reliability.
  • Participate in on-call rotations supporting production systems, conduct blameless post-mortems with root cause analysis, and maintain incident response runbooks and procedures.
  • Create comprehensive technical documentation - runbooks, architectural diagrams, network topology maps, rack elevations, and capacity models - and maintain knowledge bases for operational procedures and best practices.

Experience We're Looking For:
  • 8+ years of experience as an SRE or in a related field, with a strong focus on production systems, containers, microservices, and service delivery, required. A bachelor's degree in Computer Science, Engineering, or a related field required; an advanced degree is strongly preferred.
  • Must have extensive on-premises data center experience managing large environments with hundreds of bare-metal servers and virtual machines. Experience with multi-site or colocation deployments, cross-site disaster recovery, and hybrid on-premises and cloud trade-offs is a plus.
  • Hands-on server hardware experience is essential, including firmware and BIOS management, out-of-band management, failure diagnosis, vendor coordination, and capacity and refresh planning.
  • Deep knowledge of Linux operating systems-including configuration, performance tuning, and troubleshooting-is expected.
  • A strong understanding of Linux networking concepts and protocols is necessary, along with familiarity with data center networking fundamentals such as VLANs, routing, load balancing, DNS, and DHCP.
  • Demonstrated experience with containers and orchestration technologies, particularly Kubernetes, is required, including running clusters on bare metal or on-premises virtual machines.
  • Must bring extensive experience managing and maintaining CI/CD pipelines and the supporting technologies, including GitOps workflows, Argo CD, and Helm charts.
  • Comprehensive knowledge of observability tools such as Prometheus, the ELK Stack, log collectors, and Grafana is key to success in this role.
  • Experience with configuration management tools-particularly Ansible-is important. Infrastructure as code applied to physical infrastructure using Terraform providers, MAAS, Foreman, or similar technologies, as well as scripting in Python, Bash, or Go, would be advantageous.
  • Proven ability to analyze complex systems, identify bottlenecks, troubleshoot issues, and implement effective solutions is critical. Experience operating Kafka or other high-throughput, stateful distributed systems on premises-including replication and disaster recovery-would be highly valued.
  • Excellent communication and cross-functional collaboration skills are essential, as is a willingness to participate in an on-call rotation and support work in physical data center environments when needed. #LI-CC1


What We Offer:

Working at Ooma means being a team player, while allowing your individual voice to come through. And, you'll receive competitive compensation, benefits and generous company perks.

  • Comprehensive Medical/Dental/Vision insurance for you and eligible dependents
    • HMO, PPO's or a PPO with a HDHP (including HSA, which Ooma helps fund)
  • Employer Paid Income Protection Benefits (Basic Life and AD&D, Short- and Long-term disability)
  • FSA Healthcare & Dependent Care
  • Commuter Benefits
  • Voluntary Accident, Critical Illness, Hospital Indemnity and Legal
  • 401(k), including employer match, and Roth
  • Employee Stock Purchase Plan (ESPP)
  • Paid Time off, Sick Time, as well as corporate holidays observed
  • Employee Assistance Program
  • Life Balance benefits with Travel Assistance Services and Identity Theft
  • Additional Benefits include a Discount Program, Credit Union, Medicare Assistance, etc.


The base salary range for candidates within the United States is listed below. Actual base pay will depend on a variety of factors such as education, skills, experience, specific location, etc. The base pay range is subject to change and may be modified in the future. Regular employees may also be eligible for bonus(es), sales incentive(s) (target included in OTE) and/or stock in the form of Restricted Stock Units (RSUs).

United States Pay Range

$120,000-$170,000 USD

About Ooma

Ooma, Inc. is a telecommunications company that provides voice-over-IP (VoIP) services. Ooma's consumer home phone service has been ranked #1 in the United States for the past 10 years by Consumer Reports. The company also offers small business phone systems and home security cameras. Ooma was founded in 2004 and is headquartered in Sunnyvale, California.
Learn more about Ooma
Size
383 employees
Market Cap
$327.1 million
Industry
Net Income
-$4.1 million
Founded
2004
5 Year Trend
+13%
Revenue
$165.3 million
NASDAQ

Similar Jobs

More Jobs at Ooma

More Information Technology Jobs

Find similar Site Reliability Engineer jobs: