Role Overview:This role is for a Data Engineer with deep expertise in designing, implementing, and optimizing scalable, resilient data pipelines. The position focuses on leveraging Snowflake features, DBT models, Python ingestion frameworks, and managing database replication with Qlik Replicate. Responsibilities also include ensuring robust data governance, security, and orchestration using Astronomer Airflow and integrating with CI/CD practices across various data formats and sources.
Key Responsibilities:- Design and implement scalable, resilient data pipelines using Snowflake features including Snowpipe, Tasks, Streams, Dynamic Tables, and advanced SQL.
- Build and maintain DBT models with strong testing, documentation, and lineage.
- Develop Python ingestion frameworks for files and APIs, including schema validation, retries, and metadata capture.
- Engineer ingestion for CSV, fixed width multi record layouts, JSON, XML, Excel, and semi structured formats.
- Design Mainframe VSAM data ingestion pattern for complex EBCDIC data formats.
- Detect, analyze, and manage schema drift across file, API, and replicated database sources.
- Implement metadata driven schema evolution strategies to ensure downstream stability.
- Coordinate schema changes through controlled CI/CD workflows.
- Configure and manage Qlik Replicate tasks for CDC and full load replication from Oracle, SQL Server, and DB2.
- Ensure idempotent, auditable, and recoverable replication pipelines with strong monitoring and reconciliation.
- Implement and maintain Snowflake Data Masking policies, including dynamic masking, conditional masking, and role based masking rules.
- Apply Protegrity tokenization for sensitive data fields across ingestion and transformation layers.
- Enforce RBAC, data access controls, and governance standards across Snowflake and supporting systems.
- Build and schedule workflows using Astronomer Airflow, ensuring dependency management, retries, SLAs, and observability.
- Integrate pipelines with enterprise DevOps processes using GitLab and Azure DevOps for CI/CD automation.
- Manage code repositories using GitLab, including branching strategies, merge requests, code reviews, and approvals.
- Implement monitoring and alerting for ingestion pipelines, schema drift, replication, and transformation workloads.
- Optimize Snowflake compute, storage, and query performance; scale ingestion pipelines to meet evolving data volume and latency requirements.
Required Skills:- Deep expertise with Snowflake, including data masking policies, RBAC, performance tuning, and advanced SQL.
- Strong experience with Qlik Replicate for CDC and database replication.
- Excellent proficiency in Python and Pyspark for ingestion frameworks and automation.
- Hands-on experience with DBT Cloud and Astronomer Airflow.
- Experience with schema drift detection and schema evolution patterns.
- Experience with GitLab and CI/CD pipelines.
- Snowflake Data Build Tool (DBT).
Qualifications:- 8+ years of hands-on data engineering experience.
Preferred Skills:- Familiarity with Protegrity or similar data protection platforms.