Instrumental

Site Reliability Engineer

Instrumental$140K — $165K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3-4 years in Site Reliability Engineering, DevOps, or related roles for SaaS environments.
  • Hands-on AWS experience including EC2, VPC, IAM, RDS, ECS, and S3.
  • Infrastructure management experience using Terraform or similar technologies.
  • Design and support of CI/CD pipelines with GitHub Actions, Jenkins, or GitLab CI/CD.
  • Proficient with monitoring tools like Datadog, including dashboards and logging.
  • Experience with Docker and Kubernetes.
  • Scripting skills in Python and/or Bash.

Responsibilities

  • Operate and improve AWS-based SaaS platform.
  • Focus on reliability, automation, observability, and operational excellence.
  • Participate in a bi-weekly on-call rotation.
  • Engineer operational solutions to reduce complexity.
  • Collaborate with software engineers to enhance service readiness.
  • Continuously seek opportunities to automate repetitive tasks.
  • Support production environments including incident response.

Benefits

  • Health, vision, and dental insurance.
  • Parental leave policies.
  • Commuter plans.
Full Job Description
As a Site Reliability Engineer, you'll operate, improve, and scale our AWS-based SaaS platform. You'll combine hands-on production operations with engineering, focusing on reliability, automation, observability, and operational excellence. You'll participate in a bi-weekly on-call rotation, but the goal isn't simply to keep systems running-it's to continuously engineer away the operational complexity that comes with scaling our platform and customer base.

Requirements:

  • 3-4 years of experience in Site Reliability Engineering, DevOps, Cloud Operations, Platform Engineering, or Systems Engineering supporting production SaaS environments.
  • Strong hands-on experience with AWS, including EC2, VPC, IAM, RDS, ECS, and S3.
  • Experience managing infrastructure using Terraform or other Infrastructure as Code technologies.
  • Experience designing and supporting CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or similar platforms.
  • Strong experience with monitoring and observability tools, preferably Datadog, including dashboards, alerting, logging, and APM.
  • Experience with Docker and Kubernetes.
  • Scripting experience with Python and/or Bash.
  • Experience supporting production environments through an on-call rotation, including incident response and root cause analysis.
  • Proven ability to take ownership of production issues and drive them through investigation, remediation, and long-term resolution.


Who You Are:

  • Dead serious about performance, scalability, and reliability (PSR): You care deeply about how systems behave in the real world and continuously look for ways to make them more reliable, scalable, observable, and supportable.
  • Automation, automation, automation: If something is repetitive, manual, or error-prone, your first instinct is to automate it and make it disappear.
  • An engineer at heart: You don't want to repeatedly fight the same fires. You look for the underlying cause and build durable engineering solutions that reduce operational toil and technical debt.
  • Strong systems thinker: You understand how infrastructure, applications, networks, deployments, monitoring, and people interact-and can troubleshoot complex production issues across those boundaries.
  • Collaborative and reliable: You partner closely with software engineers to make services production-ready, improve operational workflows, and build reliability into systems before they become problems.
  • Comfortable with growth and ambiguity: You're comfortable making good decisions without perfect information and adapting as the platform, customer base, and company scale quickly.


Nice to Have:

  • Experience working in a high-growth B2B SaaS environment.
  • Experience implementing SRE practices such as SLIs, SLOs, and error budgets.
  • Experience building internal tooling and automation to eliminate operational toil.
  • Experience supporting multi-region AWS environments.
  • AWS cost optimization or FinOps experience.
  • Network, application security, and compliance experience.
  • Experience introducing AI tools or processes into engineering and operational workflows.


This position requires access to items and data that are developed under U.S. government contracts and subject to dissemination controls that limit access to U.S. citizens only.

We're a growing team that works collaboratively, is supportive of each other, and is highly energized by the opportunity for a large impact. We actively work to promote an inclusive environment, valuing passion and the ability to learn. You're encouraged to apply even if your experience doesn't precisely match the job description!

The following is a representative annual base salary range for this position within the Bay Area: $140,000-$165,000. Job level and salary opportunities are evaluated through our interview process - we review the experience, knowledge, skills, and abilities of each applicant.

Instrumental is proud to offer a highly-rated variety of benefits, including health, vision, dental, commuter plans, and parental leave.

About Instrumental

Instrumental is a technology company that provides a platform for detecting and fixing issues in manufacturing processes. The company's platform uses machine learning algorithms to analyze data from manufacturing processes and identify potential issues before they become major problems. Instrumental's platform is used by manufacturers in a variety of industries, including electronics, automotive, and consumer goods. The company was founded in 2015 and is based in San Francisco, California.
Learn more about Instrumental
Size
10 employees
Industry
Founded
2013

Similar Jobs

More Jobs at Instrumental

More Enterprise Technology Jobs

Find similar Site Reliability Engineer jobs: