Job Summary
The Data Engineer will perform end-to-end datastore migrations from on-premises DataLake environments to AWS-hosted Lakehouse platforms as part of a migration factory team. The role will focus on refactoring and migrating data pipelines, executing large-scale data transfers, modernizing legacy SQL and Apache Spark processing logic for Snowflake and Apache Iceberg environments, and ensuring data integrity through reconciliation and validation frameworks. The engineer will collaborate with business stakeholders, data owners, application teams, infrastructure teams, and global delivery teams throughout the migration lifecycle while optimizing data workloads and following established SDLC and CI/CD practices.
Key Responsibilities
• Perform end-to-end datastore migrations from on-premises DataLake environments to AWS-hosted Lakehouse platforms.
• Refactor and migrate existing data pipelines, including data extraction, transformation, and job scheduling logic.
• Execute large-scale data transfers while ensuring data integrity, completeness, reliability, and consistency between source and target platforms.
• Translate and modernize legacy SQL and Apache Spark processing logic for Snowflake and Apache Iceberg environments.
• Analyze data usage patterns, business requirements, and downstream consumption to support reusable data product development.
• Design and build data reconciliation and validation frameworks to verify data accuracy during and after migration.
• Collaborate with business stakeholders, application teams, data owners, and technical teams to perform validation and obtain migration sign-off.
• Act as a technical liaison between migration, data engineering, application, infrastructure, and business teams throughout the migration lifecycle.
• Troubleshoot data pipeline, transformation, performance, and migration-related issues and implement appropriate technical solutions.
• Optimize data pipelines and processing workloads to improve performance, scalability, and operational efficiency.
• Apply data engineering concepts including SCD Type 2, schema evolution, partitioning, clustering, normalization and denormalization, natural and surrogate keys, and data quality controls.
• Work with structured and semi-structured data formats including JSON, Avro, and Parquet.
• Follow established SDLC, source control, testing, deployment, and CI/CD practices.
• Adapt to new data technologies, migration tools, engineering standards, and workflows as required.
• Collaborate effectively with geographically distributed and global delivery teams.
Required Qualifications
• 3+ years of hands-on software development and/or data engineering experience, including coding, development, troubleshooting, and implementation of data-related solutions.
• 3+ years of experience with SQL, including development, analysis, query troubleshooting, and performance optimization.
• 3+ years of hands-on programming experience using Python and/or Java for data processing, application development, automation, or integration.
• 3+ years of experience developing or supporting ETL/ELT data pipelines, including data extraction, transformation, loading, and pipeline troubleshooting.
• Experience developing or supporting distributed data-processing solutions using Apache Spark.
• Experience with data engineering concepts including SCD Type 2, schema evolution, partitioning, clustering, normalization versus denormalization, natural versus surrogate keys, and data quality frameworks.
• Experience with one or more data and integration technologies including Kafka, ANSI SQL, FTP, Apache Spark, Hadoop, Snowflake, Apache Iceberg, and Sybase IQ.
• Experience working with structured and semi-structured data formats including JSON, Avro, and Parquet.
• Experience working within established SDLC and CI/CD processes, including source control, testing, deployment, and release practices.
• Familiarity with containerized application environments and Kubernetes.
• Demonstrated ability to troubleshoot technical issues, communicate effectively with stakeholders, collaborate across global teams, and take ownership of assigned deliverables.
• Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical field, or equivalent work experience.
Preferred Qualifications
• Experience performing large-scale data platform or datastore migration initiatives.
• Experience migrating data workloads from on-premises environments to cloud-based platforms, particularly AWS.
• Experience working with Lakehouse architectures.
• Experience with Snowflake and Apache Iceberg-based data platforms.
• Experience working within the financial services industry.
• Experience supporting complex migration programs involving multiple business, technology, and global delivery teams.