2+ years experience in SRE, Production Engineering, or DevOps
Expert level experience in architecting and troubleshooting large-scale systems
Advanced proficiency in programming languages (e.g., Python, Golang)
Deep understanding of data structures and Linux systems internals
Extensive experience with CI/CD pipelines and infrastructure as code
Familiarity with AWS services (e.g., ECS, S3, ALB, VPC)
Knowledge of containers and orchestration using Kubernetes
Bachelor's in Computer Science or Electrical Engineering (MS preferred)
Responsibilities
Define and improve engineering best practices and deployment processes
Automate continuous integration and testing processes
Manage and maintain infrastructure for large-scale systems
Design and implement monitoring systems using Prometheus and Grafana
Optimize Linux systems for performance, reliability, and security
Oversee configuration management processes
Investigate issues and collaborate with engineers to strategize solutions
Participate in on-call rotation for system reliability
Benefits
Equal opportunity employer promoting diversity and inclusivity
Full Job Description
The Opportunity
We are seeking a talented Engineer with extensive system knowledge who wants to be part of a team to build and operate large-scale systems that enables reliable and rapid deployment with effective monitoring and resilient operations.
Your Impact
Automation! Capability! Performance! Scale!
Have a strong influence in defining our engineering best practices and deployment process
Help automate the continuous integration and testing processes to enable and scale
Manage and maintain infrastructure
Own, design and implement monitoring systems such as Prometheus and Grafana
Optimize Linux systems for performance, reliability, and security
Own configuration management process(es) and build product features as appropriate
Investigate and dig into data to find the root of a problem and strategize with our engineers on solutions
Participate in on-call rotation
We're looking for someone who
Bachelor's in Computer Science or Electrical Engineering (MS preferred)
5+ years experience in SRE/Production Engineering, and IT Systems / Enterprise IT / Systems Engineering / DevOps-internal tooling, supporting internal customers
Expert level experience architecting, developing, and troubleshooting large scale systems
Advanced level proficiency with one or more programming languages (i.e. Python, Golang)
Extensive experience with CI/CD pipelines and infrastructure as code (i.e. Terraform, Ansible)
You have a strong familiarity with AWS services (i.e. ECS, S3, ALB, VPC)
You have knowledge in containers and orchestration using Kubernetes
Experience building production quality cloud infrastructure that enables reliable and rapid deployment of large-scale systems with effective monitoring and resilient operations
You thrive working in a fast paced, startup environment
You have a proven track record taking on projects from inception to launch
Solid troubleshooting fundamentals across macOS/Windows/Linux, networking basics, and security hygiene
Able to work through ambiguous problems and coordinate with IT, Security, and Engineering stakeholders
Identity and access workflows (Okta/Entra, Google Workspace/M365), provisioning automation
Salary range
Salary Range: $185,000 to $230,000 USD per year.
This salary range represents the low and high end of the estimated salary range for this position. The actual base salary offered for the role is dependent based on several factors. Our base salary is just one component of our comprehensive total rewards package.
#LI-Hybrid
About Otter.ai
Otter.ai is an AI-powered transcription service that uses machine learning algorithms to transcribe audio and video recordings. The platform is used by businesses, journalists, and other professionals to transcribe interviews, meetings, and other recordings. Otter.ai's platform is designed to be easy to use and offers a range of features, including real-time transcription and collaboration tools.