Data Platform Architecture & Pipeline Engineering
Design & Build Develop robust, scalable, and automated Lambda/Kappa data pipelines for both batch and real-time streaming data using GCP Dataflow (Apache Beam) and Cloud Data Fusion.
Data Warehousing Optimize and manage enterprise data warehouses in Google BigQuery, ensuring efficient storage, partitioning, clustering, and query performance.
Orchestration Work with orchestration tools like Cloud Composer (Apache Airflow) to manage complex data workflows and dependencies.
CI/CD & DevOps Implement Infrastructure as Code (IaC) using tools like Terraform and maintain CI/CD pipelines for seamless data deployment.
Business Intelligence & Enablement
Semantic Modeling Collaborate with analytics teams to model data and develop robust LookML models within Looker.
Performance Tuning Optimize Looker queries and BigQuery costs by implementing caching strategies, materialized views, and efficient BI Engine configurations.
Data Governance & Quality (Preferred)
Compliance & Security Implement data governance frameworks on GCP using Dataplex, Cloud Data Loss Prevention (DLP), and IAM policies to ensure data privacy and compliance.
Data Lineage & Cataloging Establish and maintain data catalogs, metadata management, and end-to-end data lineage.
Data Quality Build automated testing frameworks to monitor data quality, anomalies, and pipeline health.
Qualifications & Skills
GCP Expertise Minimum of 3-5 years of hands-on experience building production-grade data solutions on Google Cloud Platform.
Core Tools Deep, practical knowledge of BigQuery (SQL optimization, slot management) and Dataflow (Apache Beam in Python or Java).
Programming Strong proficiency in Python, Java, or Scala, alongside expert-level SQL skills.