JOB SUMMARY
We are seeking a Site Reliability Engineer (SRE) to support enterprise-scale, mission-critical applications through automation, observability, operational excellence, and AI/ML-driven reliability initiatives. The ideal candidate will have strong experience in systems administration, automation, cloud platforms, monitoring, and incident management. This role will focus on reducing operational toil, improving system availability, implementing intelligent observability solutions, and driving adoption of AIOps practices across cloud and platform environments.
KEY RESPONSIBILITIES
• Champion Site Reliability Engineering (SRE) principles and promote automation-driven operational excellence.
• Identify opportunities to automate manual operational processes and improve system reliability.
• Design, develop, and maintain automation solutions to reduce operational overhead and increase efficiency.
• Create scripts, tools, and frameworks to automate deployment, monitoring, remediation, and operational workflows.
• Design and implement AI/ML-driven observability, anomaly detection, predictive alerting, and operational response solutions.
• Expand automation coverage across deployment, monitoring, alerting, and self-healing capabilities.
• Develop and maintain application dashboards, monitoring solutions, and alerting frameworks.
• Collaborate with engineering, product, operations, and Agile teams to improve platform reliability and availability.
• Triage, diagnose, and resolve critical production incidents and system issues.
• Support change management, release activities, and production deployments with minimal operational risk.
• Develop instrumentation, validation tools, and rollout frameworks to improve deployment success rates.
• Drive adoption of AIOps platforms and machine learning-assisted observability solutions.
• Perform capacity planning and forecasting using operational metrics and trend analysis.
• Develop and support CI/CD orchestration solutions that streamline software delivery.
• Promote GitOps practices and deployment automation strategies.
• Support cloud application deployment, configuration, migration, and operational support activities.
• Perform real-time troubleshooting of mission-critical applications and platform services.
• Provide feedback to development teams to improve application reliability and operational readiness.
• Participate in on-call support and incident response activities.
• Document operational procedures, runbooks, monitoring standards, and support processes.
• Continuously improve platform performance, scalability, resilience, and observability.
REQUIRED QUALIFICATIONS
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related discipline.
• 6-8 years of enterprise systems administration, support, or Site Reliability Engineering experience.
• 6-8 years of experience developing automation scripts and operational tooling.
• 6-8 years of experience building monitoring dashboards, alerting frameworks, and observability solutions.
• Experience working within Software Development Lifecycle (SDLC) processes and continuous improvement programs.
• Hands-on experience supporting enterprise production environments.
• Strong experience with Windows Server 2019 and 2022 administration.
• Strong Linux system administration, troubleshooting, and performance tuning experience.
• Experience supporting applications hosted on virtualized infrastructure environments.
• Experience with cloud application deployment, migration, and operational support.
• Knowledge of networking concepts including DNS, DHCP, firewalls, routing, and connectivity troubleshooting.
• Understanding of distributed systems and high-availability architectures.
• Development experience with one or more of the following:
- Python
- Java
- PowerShell
- .NET
- Bash
• Experience working with SQL, Oracle, MongoDB, or similar database technologies.
• Working knowledge of NICE Actimize or similar financial crime compliance platforms.
• Experience with messaging technologies such as Kafka, RabbitMQ, Solace, or IBM MQ.
• Hands-on experience with observability platforms such as Splunk, AppDynamics, or similar tools.
• Demonstrated experience implementing AI/ML or AIOps solutions, including anomaly detection, predictive alerting, and ML-assisted monitoring.
• Strong troubleshooting, analytical, and problem-solving skills.
• Excellent verbal and written communication skills.
• Ability to work effectively in fast-paced, highly available production environments.
PREFERRED QUALIFICATIONS
• Financial Services or Banking industry experience.
• Experience working in Agile and Scrum environments.
• Hands-on experience with AIOps platforms and intelligent observability solutions.
• Experience integrating AI/ML capabilities into operational automation and CI/CD workflows.
• Experience with CI/CD tools such as Jenkins, Harness, GitHub Actions, or similar platforms.
• Knowledge of GitOps methodologies and deployment automation best practices.
• Experience with GCP, PCF, AWS, or Azure cloud platforms.
• Experience with Kubernetes, OpenShift, or other container orchestration platforms.
• Experience supporting enterprise authentication, login, or identity management ecosystems.
• Experience building self-healing systems and automated remediation frameworks.
CERTIFICATIONS
• AWS Certified SysOps Administrator preferred.
• Google Professional Cloud DevOps Engineer preferred.
• Certified Kubernetes Administrator (CKA) preferred.
• Splunk Certification preferred.
• Relevant Cloud, DevOps, SRE, AIOps, or Observability certifications are a plus.