Job Summary:
We are seeking a Cloudera Platform Engineer with 4-6 years of experience in Big Data Platform, Cloud Operations, or Infrastructure Support roles and 3+ years of hands-on experience with the Cloudera ecosystem, including CDH/CDP. The role will be responsible for day-to-day platform operations, deployments, monitoring, troubleshooting, incident handling, and operational support across Cloudera and cloud environments. The ideal candidate will have hands-on experience with CDP services, strong monitoring and troubleshooting capabilities, and a working understanding of cloud platforms and Kubernetes.
Key Responsibilities
• Execute day-to-day platform operations for Cloudera and Big Data environments.
• Support deployments, platform monitoring, troubleshooting, and operational activities.
• Provide hands-on support for CDP services including CDE, CDW, CDF, and CAI.
• Monitor platform health, alerts, performance, and operational metrics.
• Perform issue triage, troubleshooting, and escalation for platform-related incidents.
• Handle P2/P3 production incidents and support timely resolution in accordance with operational procedures.
• Troubleshoot Kubernetes environments at the pod level, including analyzing logs and resource usage.
• Support cloud-based environments and troubleshoot issues involving IAM, storage, and networking.
• Follow established runbooks, operational procedures, and platform support processes.
Required Qualifications
• 4-6 years of experience in Big Data Platform, Cloud Operations, or Infrastructure Support roles.
• 3+ years of hands-on experience with the Cloudera ecosystem, including CDH/CDP.
• Hands-on experience with CDP services including CDE, CDW, CDF, and CAI.
• Experience with platform monitoring, alerting, issue triage, and operational support.
• Experience providing production support and handling P2/P3 incidents.
• Working knowledge of cloud platforms such as AWS, Azure, or GCP, including IAM, storage, and networking concepts.
• Strong understanding of Kubernetes concepts, including pod-level troubleshooting, log analysis, and resource usage.
• Ability to execute platform operations, support deployments, and troubleshoot production issues.
• Ability to follow runbooks and established operational procedures.