SRE Product Support Manager

SysMind Tech

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Site Reliability Engineering or related fields.
  • Proficiency in monitoring tools such as Prometheus and Grafana.
  • Experience with automation and configuration management tools like Ansible and Terraform.
  • Strong knowledge of SLOs, SLIs, and SLAs in software design.
  • Ability to conduct root-cause analysis and document solutions.

Responsibilities

  • Ensure availability, scalability, and performance of production systems.
  • Design and maintain monitoring and incident response systems.
  • Implement infrastructure automation and configuration management.
  • Collaborate with teams for SLOs, SLIs, and SLAs establishment.
  • Conduct root-cause analysis for production issues.
  • Drive continuous improvement initiatives for system resilience.
  • Optimize CI/CD pipelines for secure application deployments.

Benefits

  • Collaborative work environment with cross-functional teams.
  • Opportunities for professional development and continuous improvement.
  • Participation in enterprise-level transformation projects.
  • Exposure to a variety of monitoring and automation tools.
Full Job Description
Overview

As a Site Reliability Engineer (SRE) at SysMind, you will play a key role in maintaining the performance, scalability, and availability of enterprise-grade applications and systems. You'll work at the intersection of software engineering and operations, building automation, improving observability, and ensuring the resilience of production environments.

This position is ideal for engineers who are passionate about system optimization, root-cause analysis, and continuous improvement. You'll collaborate with cross-functional teams to diagnose complex issues, enhance deployment pipelines, and design fail-safe recovery mechanisms that keep critical systems running without interruption.

Roles & Responsibilities

  • Ensure the availability, scalability, and performance of production systems across multiple environments.
  • Design, build, and maintain monitoring, alerting, and incident response systems using tools like Prometheus, Grafana, and ELK.
  • Implement infrastructure automation and configuration management using Ansible, Terraform, or similar frameworks.
  • Collaborate with development teams to establish SLOs, SLIs, and SLAs, embedding reliability into software design.
  • Conduct root-cause analysis for production incidents, documenting solutions and preventive actions.
  • Drive continuous improvement initiatives to enhance resilience, reduce downtime, and automate repetitive tasks.
  • Optimize CI/CD pipelines for reliable and secure application deployments.


If you thrive on building stable, scalable, and high-performing systems, and want to be at the heart of enterprise transformation, fill out the form below to apply for this position.

Similar Jobs

More Jobs at SysMind Tech

  • Sr. Oracle EPM Resource
    $135K — $160K *
    San Francisco, CA 94112 (San Francisco County)
    Enterprise Technology
    In-Person
  • Product Architect
    $135K — $160K *
    Dallas, TX 75217 (Dallas County)
    Enterprise Technology
    In-Person
  • Java Spring Boot Developer
    $100K — $120K *
    Raleigh, NC 27610 (Wake County)
    Enterprise Technology
    In-Person
  • Enterprise Architect
    $125K — $150K *
    Toronto, ON M3C 0E3
    Enterprise Technology
    In-Person
  • .NET Developer/CODER
    $110K — $130K *
    Whippany, NJ 07981 (Morris County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar SRE Product Support Manager jobs: