Job SummaryWe are seeking an experienced DevOps, Kubernetes, and Site Reliability Engineer with 5+ years of relevant experience to design, automate, deploy, and support reliable application platforms and CI/CD capabilities. This role focuses on improving software delivery, platform stability, operational resilience, and production reliability.
Key Responsibilities- De sign, build, and maintain CI/CD pipelines and automated deployment solutions.
- Deploy and support applications across Kubernetes and containerized environments.
- Develop automation using Python, Linux shell scripting, YAML, and related technologies.
- Manage Git/GitHub repositories, branching strategies, pull requests, and automated workflows.
- Build and maintain Kubernetes deployment configurations and Helm charts.
- Integrate automated testing, security scanning, code-quality checks, and policy controls into CI/CD pipelines.
- Troubleshoot complex application, Kubernetes, container, infrastructure, network, and deployment issues.
- Investigate production incidents, perform root-cause analysis, and implement corrective solutions.
- Apply SRE principles to improve availability, scalability, monitoring, observability, and operational readiness.
- Collaborate with application, infrastructure, network, database, cybersecurity, and production support teams.
- Participate in after-hours production support and on-call rotation as required.
Required Qualifications- Min imum 5 years of experience in DevOps, SRE, Platform Engineering, Production Engineering, Infrastructure Engineering, or a related field.
- Strong hands-on experience with Kubernetes, Docker/Podman, Linux/UNIX, Python, Bash/KornShell, Git, and GitHub.
- Experience with CI/CD and artifact-management tools such as Jenkins and Artifactory, or equivalent platforms.
- Experience with YAML, Ansible, or similar automation frameworks.
- Strong knowledge of networking concepts including DNS, TCP/IP, HTTP/HTTPS, TLS, proxies, firewalls, routing, and load balancing.
- Experience supporting production applications, troubleshooting incidents, and performing root-cause analysis.
- Understanding of SRE practices, monitoring, incident response, release automation, and deployment strategies.
- Experience with automated testing, code-quality tools, security scanning, and CI/CD policy controls.
- Experience working in Agile environments and using tools such as Jira.
- Strong communication, collaboration, troubleshooting, and prioritization skills.
Preferred Qualifications- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline, or equivalent practical industry experience.
- Experience with OpenShift or other managed Kubernetes platforms.
- Experience with Azure, AWS, GCP, infrastructure-as-code, Helm, and GitOps.
- Knowledge of service mesh, container networking, ingress controllers, and API gateways.
- Experience with observability and telemetry platforms, including Grafana.
- Knowledge of high availability, disaster recovery, capacity management, and production resiliency.
- Experience with relational databases such as DB2, Sybase, or Oracle.
- Experience within financial services or another regulated enterprise environment.
- Familiarity with secure software supply-chain practices, secrets management, certificate management, and vulnerability remediation.