Senior Site Reliability Engineer

Compunnel

$120K — $145K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • 6-8 years of experience in enterprise systems administration or Site Reliability Engineering.
  • Strong experience in building monitoring solutions and observability tools.
  • Hands-on experience with cloud application deployment and operational support.
  • Experience with AI/ML solutions in operational contexts.

Responsibilities

  • Promote automation for operational excellence in Site Reliability Engineering.
  • Automate manual operational processes to enhance system reliability.
  • Develop and maintain automation solutions that improve efficiency.
  • Implement AI/ML-driven observability and predictive alerting systems.
  • Collaborate with various teams to boost platform reliability and availability.
  • Conduct real-time troubleshooting of critical applications and services.
  • Support CI/CD orchestration to streamline software delivery.

Benefits

  • Participation in on-call support and incident response
  • Engagement in continuous improvement programs
  • Opportunity to work with cutting-edge AI/ML technologies
  • Collaboration in cross-functional teams to promote organizational growth
  • Exposure to diverse cloud environments and technologies
Full Job Description
JOB SUMMARY

We are seeking a Site Reliability Engineer (SRE) to support enterprise-scale, mission-critical applications through automation, observability, operational excellence, and AI/ML-driven reliability initiatives. The ideal candidate will have strong experience in systems administration, automation, cloud platforms, monitoring, and incident management. This role will focus on reducing operational toil, improving system availability, implementing intelligent observability solutions, and driving adoption of AIOps practices across cloud and platform environments.

KEY RESPONSIBILITIES
• Champion Site Reliability Engineering (SRE) principles and promote automation-driven operational excellence.
• Identify opportunities to automate manual operational processes and improve system reliability.
• Design, develop, and maintain automation solutions to reduce operational overhead and increase efficiency.
• Create scripts, tools, and frameworks to automate deployment, monitoring, remediation, and operational workflows.
• Design and implement AI/ML-driven observability, anomaly detection, predictive alerting, and operational response solutions.
• Expand automation coverage across deployment, monitoring, alerting, and self-healing capabilities.
• Develop and maintain application dashboards, monitoring solutions, and alerting frameworks.
• Collaborate with engineering, product, operations, and Agile teams to improve platform reliability and availability.
• Triage, diagnose, and resolve critical production incidents and system issues.
• Support change management, release activities, and production deployments with minimal operational risk.
• Develop instrumentation, validation tools, and rollout frameworks to improve deployment success rates.
• Drive adoption of AIOps platforms and machine learning-assisted observability solutions.
• Perform capacity planning and forecasting using operational metrics and trend analysis.
• Develop and support CI/CD orchestration solutions that streamline software delivery.
• Promote GitOps practices and deployment automation strategies.
• Support cloud application deployment, configuration, migration, and operational support activities.
• Perform real-time troubleshooting of mission-critical applications and platform services.
• Provide feedback to development teams to improve application reliability and operational readiness.
• Participate in on-call support and incident response activities.
• Document operational procedures, runbooks, monitoring standards, and support processes.
• Continuously improve platform performance, scalability, resilience, and observability.

REQUIRED QUALIFICATIONS
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
• 6-8 years of enterprise systems administration, support, or Site Reliability Engineering experience.
• 6-8 years of experience developing automation scripts and operational tooling.
• 6-8 years of experience building monitoring dashboards, alerting frameworks, and observability solutions.
• Experience working within Software Development Lifecycle (SDLC) processes and continuous improvement programs.
• Hands-on experience supporting enterprise production environments.
• Strong experience with Windows Server 2019 and 2022 administration.
• Strong Linux system administration, troubleshooting, and performance tuning experience.
• Experience supporting applications hosted on virtualized infrastructure environments.
• Experience with cloud application deployment, migration, and operational support.
• Knowledge of networking concepts including DNS, DHCP, firewalls, routing, and connectivity troubleshooting.
• Understanding of distributed systems and high-availability architectures.
• Development experience with one or more of the following:

- Python

- Java

- PowerShell

- .NET

- Bash
• Experience working with SQL, Oracle, MongoDB, or similar database technologies.
• Working knowledge of NICE Actimize or similar financial crime compliance platforms.
• Experience with messaging technologies such as Kafka, RabbitMQ, Solace, or IBM MQ.
• Hands-on experience with observability platforms such as Splunk, AppDynamics, or similar tools.
• Demonstrated experience implementing AI/ML or AIOps solutions, including anomaly detection, predictive alerting, and ML-assisted monitoring.
• Strong troubleshooting, analytical, and problem-solving skills.
• Excellent verbal and written communication skills.
• Ability to work effectively in fast-paced, highly available production environments.

PREFERRED QUALIFICATIONS
• Financial Services or Banking industry experience.
• Experience working in Agile and Scrum environments.
• Hands-on experience with AIOps platforms and intelligent observability solutions.
• Experience integrating AI/ML capabilities into operational automation and CI/CD workflows.
• Experience with CI/CD tools such as Jenkins, Harness, GitHub Actions, or similar platforms.
• Knowledge of GitOps methodologies and deployment automation best practices.
• Experience with GCP, PCF, AWS, or Azure cloud platforms.
• Experience with Kubernetes, OpenShift, or other container orchestration platforms.
• Experience supporting enterprise authentication, login, or identity management ecosystems.
• Experience building self-healing systems and automated remediation frameworks.

CERTIFICATIONS
• AWS Certified SysOps Administrator preferred.
• Google Professional Cloud DevOps Engineer preferred.
• Certified Kubernetes Administrator (CKA) preferred.
• Splunk Certification preferred.
• Relevant Cloud, DevOps, SRE, AIOps, or Observability certifications are a plus.

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Senior Site Reliability Engineer jobs: