Sonar

Major Incident Manager

Sonar$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Professional experience with Infrastructure as Code (IaC) using Terraform or similar tools
  • Hands-on experience with a major cloud provider (AWS, GCP, Azure)
  • Practical experience with Service Level Indicators (SLIs) and Service Level Objectives (SLOs)
  • Experience with modern observability platforms like Prometheus or Datadog
  • Strong understanding of core networking concepts
  • Experience implementing security best practices via code
  • Practical experience managing large-scale Identity and Access Management systems

Responsibilities

  • Monitor critical security infrastructure and maintain dashboards for Service Level Objectives (SLOs)
  • Develop and maintain Infrastructure as Code for automated deployment and configuration
  • Identify and implement automation solutions for manual operational tasks
  • Enhance the DevSecOps security tools within CI/CD pipelines
  • Participate in on-call rotations and develop engineering solutions from incident post-mortems

Benefits

  • Relocation support for candidates willing to move to the job location
  • Opportunity to work collaboratively in a hybrid office environment with designated in-office days
  • Involvement in a growing company with ongoing development of tools and processes
  • Work in a team that supports diverse international offices
Full Job Description
Position description

We are still at the beginning of our growth journey, so we are putting new processes, technologies, and tools in place on a continuous basis. Your role is a pivotal engineering contributor to the tooling and services to automate and enhance the software development lifecycle, empowering our fellow SonarSourcers to deliver with speed, confidence, and security. You would be a member of a team that delivers solutions across all of our 5 offices: Austin (Texas, US), Geneva (Switzerland), Bochum (Germany) and Singapore.

As a Major Incident Manager, you use and create automation tools to monitor and observe production infrastructure services both on premises and in the cloud. You are allergic to repetitive tasks, preferring to maximize automation and reliability. You are expert in change management, infrastructure management, system support, and configuration management.

What you will do

  • System Health Monitoring, Alert Triaging, and Error Budget Management: Dedicate time to monitoring critical security infrastructure (e.g., identity platforms, firewalls, compliance systems) and core infrastructure components. Focus on using and maintaining dashboards tied to Service Level Objectives (SLOs), triaging high-severity alerts, and analyzing the current Error Budget burn rate to guide prioritization for the rest of the day.
  • Infrastructure as Code (IaC) and Policy as Code Development: Spend the largest portion of time writing, reviewing, and testing code (e.g., Python, Go, Terraform, or proprietary tools) to automate the deployment, configuration, and security hardening of infrastructure. This involves treating infrastructure and security policies as software to ensure consistency and prevent configuration drift.
  • Toil Elimination and Automation of Operational Tasks: Identify, scope, and implement automated solutions for manual, repetitive, and time-consuming tasks (toil) related to security patching, compliance checks, certificate rotations, or infrastructure maintenance. The goal is to continuously reduce the operational workload for the team.
  • Security Pipeline and Observability Maintenance: Maintain and enhance the DevSecOps security tools integrated into the CI/CD pipelines (e.g., static analysis, vulnerability scanning, security configuration checks). Ensure the end-to-end logging, metrics, and tracing (observability) systems for both infrastructure and security tools are robust, accurate, and provide immediate diagnostic capability during incidents.
  • Incident Response Engineering and Post-Mortem Action: Participate in the on-call rotation and actively engage in engineering solutions derived from post-mortems. This means turning incident root causes into preventative measures implemented via code, improving runbooks into automated actions, and reducing Mean Time To Resolution (MTTR) for future incidents.


Experience and qualifications

  • Deep IaC Expertise: Professional experience provisioning and managing complex infrastructure using tools like Terraform or CloudFormation (AWS), or similar tools like Ansible or Puppet for configuration management.
  • Cloud/Platform Experience: Hands-on experience with a major cloud provider (AWS, GCP, Azure) or managing large-scale internal/private cloud infrastructure.
  • SLO/SLI Implementation: Practical experience defining, measuring, and reporting on Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for critical services.
  • Logging/Metrics/Tracing Stacks: Proven experience with modern observability platforms (e.g., Prometheus/Grafana, ELK/EFK stack, proprietary systems, or vendor solutions like Datadog/Splunk) for proactive issue identification.
  • Networking: Strong understanding of core networking concepts (TCP/IP, DNS, Load Balancing, Firewalls, Proxies) sufficient to debug complex service connectivity and latency issues.
  • Automation of Security Controls: Experience implementing security best practices via code, such as automated vulnerability scanning, configuration hardening, secret management (e.g., HashiCorp Vault), and key rotation.
  • Identity and Access Management (IAM): Practical experience managing large-scale IAM systems (e.g., implementing least-privilege policies, single sign-on).
  • Incident Management: Experience running or significantly contributing to post-incident reviews (post-mortems) and prioritizing resulting engineering work (error budget management).


In-office culture

We're intentional about this. We believe the best teams are built in the room together. Three anchor days - Mondays, Tuesdays, and Thursdays - create the collaboration rhythm that makes a hub office worth having.

Candidates need to be genuinely based in the location the role is posted - if that's not where you are today, we're happy to support relocation for the right person.

About Sonar

Sonar is a technology company that provides a platform for businesses to manage their customer feedback. The company was founded in 2017 and has since been helping businesses to collect, analyze, and act on customer feedback. Sonar's platform uses artificial intelligence and natural language processing to analyze customer feedback and provide insights to businesses. The company has a team of experienced software developers, data scientists, and customer success managers who work together to deliver high-quality services to clients. Sonar's platform is used by businesses in various industries, including retail, hospitality, and healthcare.
Learn more about Sonar
Size
50 employees
Industry

Similar Jobs

More Jobs at Sonar

More Information Technology Jobs

Find similar Major Incident Manager jobs: