Site Reliability Engineer (SRE)Job Title Site Reliability Engineer (SRE)
Job Summary We are seeking a skilled Site Reliability Engineer (SRE) to build, automate, and maintain highly available, scalable, and reliable infrastructure and applications. The ideal candidate will combine software engineering and operations expertise to improve system reliability, automate operational processes, optimize performance, and ensure service availability. You will work closely with software developers, DevOps engineers, cloud architects, and security teams to support production systems and drive operational excellence.
Key Responsibilities - Design, implement, and maintain highly available and scalable infrastructure for production environments.
- Automate infrastructure provisioning, deployment, monitoring, and operational workflows.
- Improve system reliability, availability, scalability, and performance through engineering best practices.
- Develop and maintain CI/CD pipelines to support automated application deployments.
- Monitor production systems, troubleshoot incidents, perform root cause analysis (RCA), and implement preventive measures.
- Configure and manage observability solutions including logging, monitoring, tracing, and alerting.
- Collaborate with development teams to improve application reliability and operational readiness.
- Implement disaster recovery, backup, and business continuity strategies.
- Optimize cloud infrastructure utilization and operational costs.
- Develop automation scripts and tools to eliminate repetitive operational tasks.
- Participate in incident response, on-call rotations, and post-incident reviews.
- Document infrastructure architecture, operational procedures, and system configurations.
Required Qualifications - Bachelor's degree in Computer Science, Information Technology, Software Engineering, or a related field.
- 3-6 years of experience in Site Reliability Engineering, DevOps, Cloud Engineering, or Systems Engineering.
- Strong proficiency in Linux system administration.
- Experience with Python, Go, Bash, or Shell scripting.
- Hands-on experience with Docker and Kubernetes.
- Strong understanding of networking concepts including TCP/IP, DNS, HTTP/HTTPS, and load balancing.
- Experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI/CD, or Azure DevOps.
- Familiarity with cloud platforms such as AWS, Microsoft Azure, or Google Cloud.
- Strong understanding of infrastructure automation and configuration management.
Preferred Qualifications - Experience with Infrastructure as Code (IaC) tools such as Terraform, Ansible, Pulumi, or CloudFormation.
- Hands-on experience with monitoring and observability tools including Prometheus, Grafana, ELK Stack, OpenTelemetry, Datadog, or Splunk.
- Familiarity with service mesh technologies such as Istio or Linkerd.
- Experience managing distributed systems and microservices architectures.
- Knowledge of security best practices, identity management, and compliance requirements.
- Experience supporting high-availability, mission-critical production systems.
Technical Skills - Linux
- Python
- Go
- Bash / Shell Scripting
- Docker
- Kubernetes
- Terraform
- Ansible
- Jenkins
- GitHub Actions
- GitLab CI/CD
- Prometheus
- Grafana
- ELK Stack
- OpenTelemetry
- Datadog
- Splunk
- Git
- AWS / Azure / Google Cloud
- Networking (TCP/IP, DNS, HTTP/HTTPS)
- SQL (basic)
Soft Skills - Strong analytical and troubleshooting abilities.
- Excellent problem-solving and incident management skills.
- Effective communication and collaboration across engineering teams.
- Ability to perform under pressure during production incidents.
- Strong documentation and organizational skills.
- Continuous improvement mindset and commitment to automation.
Nice to Have - Experience supporting AI/ML infrastructure or MLOps platforms.
- Familiarity with container security and Kubernetes security best practices.
- Knowledge of FinOps and cloud cost optimization.
- Experience with chaos engineering and resilience testing.
- Relevant certifications such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), Google Professional Cloud DevOps Engineer, or Microsoft Azure DevOps Engineer Expert.
Benefits - Competitive salary and performance-based incentives.
- Comprehensive health and wellness benefits.
- Flexible or hybrid work arrangements.
- Learning, certification, and conference sponsorship opportunities.
- Access to modern cloud infrastructure and engineering tools.
- Opportunity to work on highly scalable, mission-critical systems in a collaborative and innovative environment.