Role Overview:This role is for a Data Engineer focused on designing, implementing, and optimizing robust ETL/ELT pipelines and complex data models primarily using PySpark and Snowflake. The engineer will be responsible for ensuring high data quality, integrity, and efficient operations within the data infrastructure.
Key Responsibilities:- Design and implement robust ETL/ELT pipelines using PySpark and SQL to ingest data from diverse sources including APIs, flat files, and relational databases.
- Develop and manage complex data models (e.g., Star/Snowflake schemas) and maintain Fact and Dimension tables within Snowflake.
- Monitor and tune Snowflake queries and Spark jobs to optimize performance, reduce latency, and manage computational costs.
- Implement automated data validation frameworks and testing procedures to ensure a "single source of truth" and high data reliability.
- Troubleshoot production pipeline issues and manage version control via Git.
Required Skills:- Snowflake
- PySpark
- SQL
- ETL/ELT pipeline development
- Data modeling (Star/Snowflake schemas)
- Data validation frameworks
- Git for version control
Qualifications:- 6-8 years of experience in data engineering.