Data Platform Architecture & Pipeline Engineering
Design & Build: Develop robust, scalable, and automated Lambda/Kappa data pipelines for both batch and real-time streaming data using GCP Dataflow (Apache Beam) and Cloud Data Fusion.
Data Warehousing: Optimize and manage enterprise data warehouses in Google BigQuery, ensuring efficient storage, partitioning, clustering, and query performance.
Orchestration: Work with orchestration tools like Cloud Composer (Apache Airflow) to manage complex data workflows and dependencies.
CI/CD & DevOps: Implement Infrastructure as Code (IaC) using tools like Terraform and maintain CI/CD pipelines for seamless data deployment.
Business Intelligence & Enablement
Semantic Modeling: Collaborate with analytics teams to model data and develop robust LookML models within Looker.
Performance Tuning: Optimize Looker queries and BigQuery costs by implementing caching strategies, materialized views, and efficient BI Engine configurations.
Data Governance & Quality (Preferred)
Compliance & Security: Implement data governance frameworks on GCP using Dataplex, Cloud Data Loss Prevention (DLP), and IAM policies to ensure data privacy and compliance.
Data Lineage & Cataloging: Establish and maintain data catalogs, metadata management, and end-to-end data lineage.
Data Quality: Build automated testing frameworks to monitor data quality, anomalies, and pipeline health.
Qualifications & Skills
GCP Expertise: Minimum of 3-5 years of hands-on experience building production-grade data solutions on Google Cloud Platform.
Core Tools: Deep, practical knowledge of BigQuery (SQL optimization, slot management) and Dataflow (Apache Beam in Python or Java).
Programming: Strong proficiency in Python, Java, or Scala, alongside expert-level SQL skills.