Sr SRE Automation Engineer

Compunnel

$120K — $145K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, IT, Engineering, or equivalent experience.
  • Experience with Site Reliability Engineering in enterprise settings.
  • Proficient in building automation frameworks and operational tooling.
  • Strong scripting skills in Python, PowerShell, Bash, Java, or .NET.
  • Expertise in cloud platforms, distributed systems, and high-availability architectures.
  • Skilled in designing monitoring, logging, and alerting solutions.
  • Familiarity with AI/ML operational automation and predictive monitoring.

Responsibilities

  • Champion SRE principles and promote automation-driven excellence.
  • Develop innovative tools to solve complex operational challenges.
  • Implement automation solutions to enhance operational efficiency.
  • Create scripts and frameworks to streamline infrastructure management.
  • Design AI/ML-driven monitoring and operational response solutions.
  • Lead efforts to expand automation across various operational workflows.
  • Collaborate with cross-functional teams to improve system reliability.

Benefits

  • Comprehensive health coverage including medical, dental, and vision.
  • Flexible work hours and remote work options.
  • Professional development opportunities and training.
  • Access to advanced tools and technologies in a cutting-edge environment.
  • Emphasis on work-life balance with supportive company culture.
Full Job Description
JOB SUMMARY

We are seeking a Site Reliability Engineer (SRE) to drive automation, reliability, observability, and operational excellence across enterprise-scale, mission-critical applications. The ideal candidate will have strong experience in automation, cloud platforms, AIOps, monitoring, incident management, and production support. This role focuses on reducing operational toil, improving system availability, implementing AI/ML-driven operational solutions, and advancing platform reliability through modern SRE practices.

KEY RESPONSIBILITIES
• Champion Site Reliability Engineering (SRE) principles and promote automation-driven operational excellence.
• Identify opportunities to build innovative tools and solutions that address complex operational challenges across enterprise and mission-critical applications.
• Design, develop, and maintain automation solutions that reduce manual effort and improve operational efficiency.
• Create scripts and automation frameworks to streamline infrastructure management, deployment processes, and operational workflows.
• Design and implement AI/ML-driven automation pipelines, anomaly detection, predictive alerting, and intelligent operational response solutions.
• Enhance observability capabilities through advanced monitoring, telemetry, logging, and analytics platforms.
• Lead the expansion of automation coverage across deployment, monitoring, alerting, remediation, and self-healing workflows.
• Collaborate with Engineering, Scrum, Operations, and Infrastructure teams to improve system availability, reliability, and performance.
• Monitor, triage, troubleshoot, and resolve critical production incidents and platform issues.
• Implement operational changes with minimal risk while ensuring effective stakeholder communication.
• Develop tools, frameworks, dashboards, and instrumentation to improve application deployment success and operational visibility.
• Leverage AI/ML capabilities to enhance platform monitoring, rollout validation, and operational intelligence.
• Drive adoption of AIOps platforms and machine learning-assisted observability practices.
• Support capacity planning and performance forecasting using data-driven analytics and predictive models.
• Design and implement CI/CD orchestration solutions to accelerate software delivery and improve deployment reliability.
• Promote GitOps methodologies and automation best practices across engineering teams.
• Troubleshoot mission-critical application workflows and collaborate with development teams to address reliability concerns.
• Develop and maintain operational runbooks, support procedures, and knowledge documentation.
• Participate in on-call support rotations and incident response activities.
• Continuously identify opportunities to improve platform scalability, resilience, performance, and operational efficiency.

REQUIRED QUALIFICATIONS
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent professional experience.
• Experience implementing Site Reliability Engineering (SRE) practices within enterprise environments.
• Strong experience building automation frameworks and operational tooling.
• Experience developing scripts using Python, PowerShell, Bash, Java, .NET, or similar technologies.
• Strong knowledge of cloud platforms, distributed systems, and high-availability architectures.
• Experience designing and implementing monitoring, logging, observability, and alerting solutions.
• Knowledge of AI/ML-driven operational automation, anomaly detection, and predictive monitoring.
• Experience supporting production environments and mission-critical applications.
• Strong troubleshooting, debugging, and root cause analysis skills.
• Experience with CI/CD pipelines, deployment automation, and DevOps practices.
• Knowledge of GitOps concepts and modern software delivery methodologies.
• Strong understanding of networking concepts, infrastructure, and system administration.
• Excellent verbal and written communication skills.
• Ability to collaborate effectively across engineering, operations, and business teams.
• Strong analytical and problem-solving abilities.

PREFERRED QUALIFICATIONS
• Experience with AIOps platforms and ML-assisted observability tools.
• Experience with cloud platforms such as AWS, Azure, GCP, or PCF.
• Experience with Kubernetes, OpenShift, or container orchestration platforms.
• Experience with Splunk, AppDynamics, Grafana, Prometheus, or similar monitoring solutions.
• Knowledge of capacity planning, forecasting, and performance engineering.
• Experience supporting enterprise authentication, login, or identity management platforms.
• Experience integrating AI/ML capabilities into operational processes and CI/CD workflows.
• Financial Services or Banking industry experience.
• Experience working in Agile and Scrum environments.

CERTIFICATIONS
• Certified Kubernetes Administrator (CKA) preferred.
• AWS Certified DevOps Engineer or SysOps Administrator preferred.
• Google Professional Cloud DevOps Engineer preferred.
• Splunk Certification preferred.
• Relevant Cloud, DevOps, SRE, AIOps, or Observability certifications are a plus.

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Sr SRE Automation Engineer jobs: