Clearwater Analytics

Sr. Site Reliability Engineer

Clearwater Analytics • $120K — $145K *
Boise, ID 83709In-Person
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in computer science or related field, or equivalent experience.
  • 7+ years in Site Reliability Engineering or Platform Engineering.
  • Proven experience in incident management and root cause analysis.
  • Strong expertise in Terraform and Infrastructure as Code practices.
  • Hands-on with AWS and EKS platforms.
  • Deep understanding of monitoring/observability tools: Prometheus, Grafana, Dynatrace, OpenSearch.
  • Proficiency in programming languages like Python, Java, Go, or Bash.

Responsibilities

  • Design and maintain scalable, reliable production systems.
  • Define and manage SLIs, SLOs, and SLAs for system reliability.
  • Automate infrastructure provisioning and operations using Terraform.
  • Operate and manage AWS-based cloud-native platforms, like EKS.
  • Implement monitoring, logging, and alerting with observability tools.
  • Lead incident management, including on-call support and troubleshooting.
  • Utilize AI-assisted investigations to enhance incident response.

Benefits

  • Opportunity for professional growth in a cutting-edge technology environment.
  • Work with advanced cloud-native tools and practices.
  • Engagement in high-impact projects with significant organizational visibility.
  • Flexible working arrangements, promoting work-life balance.
Full Job Description
Job Summary

We are seeking a highly skilled Senior Site Reliability Engineer to ensure the reliability, scalability, and performance of our cloud-native systems and applications. This role drives automation, monitoring, and incident management practices while operating Kubernetes platforms (Amazon EKS) and leveraging observability tools such as Prometheus, Grafana, Dynatrace, and OpenSearch to maintain high availability and operational excellence.

Key Responsibilities
• Design, build, and maintain highly available, scalable, and reliable production systems.
• Define and manage SLIs, SLOs, and SLAs to drive system reliability.
• Automate infrastructure provisioning and operations using Terraform (IaC).
• Operate and manage cloud-native platforms, including Amazon EKS.
• Implement and maintain monitoring, logging, and alerting using Prometheus, Grafana, Dynatrace, and OpenSearch.
• Lead incident management - on-call rotation, production troubleshooting, and root cause analysis (RCA).
• Drive AI-assisted investigations as a core part of incident response, and build and maintain the prompts, integrations, and guardrails that make AI-driven triage and RCA reliable.
• Improve system reliability through automation, self-healing mechanisms, and performance tuning.
• Collaborate with development teams to improve application reliability, scalability, and deployment processes.
• Build and maintain CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions) for fast, reliable software delivery.
• Perform capacity planning and cost optimization for infrastructure and services.
• Ensure security, compliance, and best practices across infrastructure and applications.

Required Qualifications
• Bachelor's degree in computer science or a related field, or equivalent practical experience.
• 7+ years in Site Reliability Engineering or Platform Engineering.
• Proven ownership of incident management, on-call support, and root cause analysis (RCA) for production systems.
• Strong expertise in Terraform and Infrastructure as Code.
• Hands-on experience with AWS and EKS.
• Strong understanding of monitoring, logging, and observability (Prometheus, Grafana, Dynatrace, OpenSearch).
• Proficiency in Python, Java, Go, or Bash.
• Experience with Agile development and CI/CD pipelines (GitLab CI, Jenkins, or GitHub Actions).
• Strong problem-solving, documentation, and communication skills.
• Proven ability to troubleshoot effectively in high-pressure production environments.
• Experience with autoscaling, performance tuning, and cost optimization.
• Familiarity with AI-assisted automation tools and a track record of using them to reduce toil and improve reliability.

Preferred Skills
• Docker and Linux administration.
• Build systems and dependency management (Maven, Gradle, npm).
• Additional AWS services: Cognito, WAF, Elasticsearch, SNS, SQS, S3, Systems Manager.
• Database infrastructure knowledge (RDS, MySQL, SQL Server).
• Cloud or Kubernetes certifications.

About Clearwater Analytics

Clearwater Analytics is a global SaaS solution for automated investment data aggregation, reconciliation, accounting, and reporting. Clearwater helps thousands of organizations make the most of investment portfolio data with cloud-native software and client-centric servicing. Every day, investment professionals worldwide trust Clearwater to deliver timely, validated investment data and in-depth reporting. Clearwater aggregates, reconciles, and reports on more than $5.5 trillion in assets across thousands of accounts daily for our Fortune 500 clients.
Learn more about Clearwater Analytics
Size
1,500 employees
Market Cap
$4.4 billion
Industry
NASDAQ

Similar Jobs

More Jobs at Clearwater Analytics

More Information Technology Jobs

Find similar Sr. Site Reliability Engineer jobs: