5-7 years of experience in designing and developing data pipelines using AWS services.
Proficient in AWS Glue, Athena, Lambda, and Redshift.
Hands-on expertise in building ETL jobs with PySpark and SQL.
Experience with structured and semi-structured data processing.
Familiarity with data warehouse, data lake, and lakehouse architectures.
Understanding of data modeling techniques.
Ability to optimize data pipelines for performance and cost efficiency.
Exposure to CI/CD processes and data security best practices.
Responsibilities
Design and develop batch and real-time data pipelines using AWS services.
Build complex ETL and data processing jobs with PySpark and SQL.
Ingest and process structured and semi-structured data.
Support data warehouse, data lake, and lakehouse architectures.
Apply data modeling techniques for analytics requirements.
Optimize data pipelines for performance and cost efficiency.
Participate in CI/CD processes for data engineering solutions.
Implement best practices for securing data pipelines and assets.
Benefits
Opportunities for professional development and training.
Collaborative work environment with a focus on innovation.
Access to cutting-edge AWS technologies and tools.
Flexible work arrangements to support work-life balance.
Full Job Description
Job Summary The AWS Data Engineer will design and develop batch and real-time data pipelines using AWS data and analytics services. The role will focus on building complex ETL and data processing solutions using PySpark and SQL, working with structured and semi-structured data, and supporting data warehouse, data lake, and lakehouse architectures. The position will also contribute to pipeline optimization, CI/CD, data security, quality, and governance.
Key Responsibilities • Design and develop batch and real-time data pipelines using AWS data and analytics services such as AWS Glue, Athena, Lambda, and Redshift. • Build complex data processing and ETL jobs using PySpark and SQL. • Ingest and process structured and semi-structured data. • Support data warehouse, data lake, and lakehouse architectures. • Apply appropriate data modeling techniques to support data processing and analytics requirements. • Optimize data pipelines to improve performance and cost efficiency. • Participate in CI/CD processes for data engineering solutions. • Apply best practices for securing data pipelines and data assets within AWS. • Support data quality and data governance processes where required.
Required Qualifications • Experience designing and developing batch and real-time data pipelines using AWS data and analytics services. • Strong experience with AWS Glue, Athena, Lambda, and Redshift. • Hands-on experience building complex ETL and data processing jobs using PySpark and SQL. • Experience ingesting and processing structured and semi-structured data. • Understanding of data warehouse, data lake, and lakehouse architectures. • Good understanding of data modeling techniques. • Ability to optimize data pipelines for performance and cost efficiency. • Exposure to CI/CD processes. • Knowledge of best practices for securing AWS data pipelines and data assets.
Preferred Qualifications • Knowledge of data quality processes. • Knowledge of data governance processes.