JOB SUMMARY
The Senior Platform Engineer / Site Reliability Engineer (SRE) will be responsible for designing, implementing, and optimizing enterprise messaging platforms, automation frameworks, observability solutions, and DevOps capabilities that support mission-critical applications. This role will drive platform modernization, operational excellence, reliability engineering, and automation initiatives while improving developer productivity and reducing operational overhead. The ideal candidate will possess strong expertise in RabbitMQ, Kafka, Python or Ansible automation, monitoring platforms, and enterprise-scale infrastructure engineering.
KEY RESPONSIBILITIES
Messaging Platform Engineering
• Design, implement, maintain, and optimize enterprise messaging platforms using RabbitMQ and Apache Kafka.
• Support highly available, fault-tolerant, and scalable messaging architectures.
• Troubleshoot messaging platform incidents, performance bottlenecks, latency issues, and capacity constraints.
• Develop standards, best practices, and operational procedures for messaging services.
• Partner with application teams to integrate messaging technologies into enterprise solutions.
Automation & Platform Engineering
• Create and maintain automation solutions using Python, Ansible, and related scripting technologies.
• Build reusable automation frameworks and self-service engineering capabilities.
• Automate provisioning, configuration management, deployment, monitoring, and operational processes.
• Support platform modernization and operational transformation initiatives.
• Improve developer productivity through platform enhancements and workflow automation.
• Develop reusable templates, scripts, and engineering standards across enterprise environments.
Observability & Reliability Engineering
• Design and implement enterprise observability and monitoring solutions.
• Configure and support monitoring platforms such as:
o Grafana
o Prometheus
o Splunk
o Dynatrace
• Establish monitoring standards, dashboards, alerting frameworks, and operational reporting.
• Monitor application health, infrastructure performance, messaging platforms, and service reliability.
• Conduct root cause analysis and implement long-term corrective actions.
• Drive improvements in platform reliability, availability, and operational visibility.
DevOps & DevSecOps
• Support CI/CD pipeline implementation, enhancement, and operational support.
• Integrate security controls and compliance checks into software delivery workflows.
• Implement and support DevSecOps tools including:
o Qualys
o Fortify
o Snyk
o SonarQube
• Partner with security teams to automate vulnerability detection, remediation, and compliance validation.
• Support software supply chain security and secure development practices.
Identity & Access Governance
• Support implementation of:
o Role-Based Access Control (RBAC)
o Single Sign-On (SSO)
o Privileged Access Management (PAM)
o Secrets Management Solutions
• Collaborate with security and infrastructure teams to establish secure access governance models.
• Support automation and lifecycle management for privileged credentials and access controls.
Database & Infrastructure Automation
• Support Database-as-Code initiatives across enterprise environments.
• Implement and maintain Liquibase-based deployment frameworks and automation processes.
• Develop automated database deployment, validation, rollback, and governance workflows.
• Support infrastructure standardization through Infrastructure-as-Code and automation practices.
REQUIRED QUALIFICATIONS
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline.
• Minimum 5 years of experience in:
o DevOps Engineering
o Platform Engineering
o Site Reliability Engineering (SRE)
o Infrastructure Automation
o Enterprise Technology Operations
• Experience supporting enterprise-scale technology environments and mission-critical applications.
• Strong hands-on expertise with:
o RabbitMQ
o Apache Kafka
• Strong scripting and automation experience using:
o Python
o Ansible
• Experience implementing and supporting observability platforms including:
o Grafana
o Prometheus
o Splunk
o Similar Enterprise Monitoring Solutions
• Experience building and supporting CI/CD pipelines and automation frameworks.
• Strong understanding of:
o DevOps Practices
o Platform Engineering
o Reliability Engineering
o Distributed Systems
• Experience collaborating with development, database, infrastructure, and security teams.
• Proven success delivering modernization, automation, or operational transformation initiatives.
• Strong troubleshooting, analytical, and root cause analysis skills.
• Excellent verbal and written communication skills.
• Strong ability to work independently and manage multiple initiatives simultaneously.
PREFERRED QUALIFICATIONS
• Experience implementing Database-as-Code frameworks, preferably Liquibase.
• Experience with Dynatrace observability and monitoring solutions.
• Experience integrating and administering:
o Qualys
o Fortify
o Snyk
o SonarQube
• Experience implementing:
o RBAC
o SSO
o PAM
o Secrets Management Solutions
• Experience designing reusable automation frameworks and self-service engineering platforms.
• Experience working in Agile, Scrum, or Kanban environments.
• Strong understanding of cloud-native architectures and platform modernization.
• Experience supporting highly available distributed applications and services.
• Experience improving developer experience and engineering productivity through platform innovations.
CERTIFICATIONS
• AWS Certified DevOps Engineer - Professional (Preferred).
• Microsoft Certified: DevOps Engineer Expert (Preferred).
• Certified Kubernetes Administrator (CKA) (Preferred).
• HashiCorp Terraform Associate (Preferred).
• Dynatrace Associate Certification (Preferred).
• Splunk Core Certified Power User or Administrator (Preferred).
• Red Hat Certified Engineer (RHCE) (Preferred).
TECHNICAL SKILLS
• RabbitMQ
• Apache Kafka
• Python
• Ansible
• Grafana
• Prometheus
• Splunk
• Dynatrace
• DevOps
• Site Reliability Engineering (SRE)
• Platform Engineering
• CI/CD
• Liquibase
• DevSecOps
• RBAC
• SSO
• Privileged Access Management (PAM)
• Secrets Management
• Qualys
• Fortify
• Snyk
• SonarQube
• Infrastructure Automation
• Monitoring & Observability
• Reliability Engineering
• Distributed Systems
• Root Cause Analysis
MANDATORY SKILLS
• RabbitMQ
• Apache Kafka
• Python or Ansible Scripting
• Grafana, Prometheus, Splunk, or Similar Observability Platforms
• DevOps / Platform Engineering
• Site Reliability Engineering (SRE)
• CI/CD Pipeline Support
• Infrastructure Automation
• Enterprise Monitoring & Alerting
• Troubleshooting and Root Cause Analysis
PREFERRED SKILLS
• Liquibase
• Dynatrace
• Qualys
• Fortify
• Snyk
• SonarQube
• RBAC
• SSO
• PAM
• Secrets Management
• Self-Service Platform Engineering
• Agile / Scrum / Kanban
• Enterprise Automation Platforms