Role Description:
Experienced GCP Data Engineer with 8-10 years of expertise in designing, developing, and optimizing large-scale data platforms and cloud-native data solutions. Strong hands-on experience in PySpark, Python, BigQuery, Dataproc, Kafka, Airflow, SQL, and GCP services, with proven success in building scalable ETL/ELT pipelines, batch and real-time data processing frameworks, and enterprise data warehouse solutions. Adept at data modeli ng, performance tuning, cloud security, automation, and collaborating with cross-functional teams to deliver high-quality data solutions that support analytics, reporting, and business intelligence initiatives.
Key Skills & Responsibilities
1. Design and develop scalable ETL/ELT pipelines using PySpark, Python, and Spark SQL.
2. Build and optimize large-scale data processing solutions on GCP using Dataproc, BigQuery, and GCS.
3. Develop and maintain Kafka-based streaming pipelines for real-time data ingestion and processing.
4. Create, manage, and optimize BigQuery datasets, tables, views, stored procedures, and complex SQL queries.
5. Implement workflow orchestration and scheduling using Cloud Composer (Airflow).
6. Manage cloud security through IAM roles, service accounts, and access control policies following best practices.
7. Work extensively with Hive, Hadoop, Spark, and distributed data processing technologies.
8. Develop automation and monitoring solutions using Python and Shell scripting to improve operational efficiency.
9. Utilize Git, CI/CD pipelines, and Terraform for code management, deployment, and infrastructure automation.
10. Collaborate with Data Architects, DevOps teams, Data Scientists, and business stakeholders while providing production support, troubleshooting, and performance optimization.