Full Job Description
We are seeking a skilled Compute Platform SRE Lead to support EPAM's Compute Managed Services project. The role focuses on KTLO (Keep the Lights On) activities, ensuring 24x7 monitoring, incident management, and operational stability across multi-cloud environments (GCP, AWS, Azure). The SRE will drive observability improvements, automate processes, and maintain compliance while collaborating with cross-functional teams to deliver high-quality compute services. Note: This role is for a Cloud Compute Platform Managed Services engagement which will require on-call support & overlap with US/UK/AU business hours as needed (On-call & shift allowance to be provided). Req.#[redacted] Responsibilities Perform 24x7 monitoring of compute platforms using tools like ELK and PagerDuty Manage incidents and problems across servers, middleware, OS, and cloud platforms, including troubleshooting, root cause analysis (RCA), and resolution Execute repaving activities, change management, and disaster recovery processes Ensure security and vulnerability compliance, including user management and certificate lifecycle management Handle service requests, configuration updates, and audit-related data extracts.Develop and update Standard Operating Procedures (SOPs) for infrastructure operations Collaborate on cell-based automation improvements and continuous service enhancements Requirements 7+ years of experience in cloud platforms AWS and GCP and OS administration (Windows/Linux) Proficiency in automation tools (Ansible, Terraform, Python, Bash) Strong knowledge of observability tools (ELK, Grafana) and incident management processes Experience in security hardening, vulnerability management, and compliance Excellent problem-solving, communication, and collaboration skills Familiarity with disaster recovery and operational recovery processes