AWS Certified (e.g. Solutions Architect, Data Analytics Specialty, Developer Associate)
Hands-on experience with AWS services: S3, EC2, EMR, Glue, Athena, Lambda, and VPCs
Strong proficiency in PySpark for data transformations and analytics
Practical experience with Apache Iceberg for data lakes management
In-depth knowledge of Apache Hive for querying and schema management
Expert-level proficiency in Python for scripting and automation
Proficient in shell scripting and SQL for data manipulation and querying
Responsibilities
Design, develop, and deploy scalable data solutions on AWS services like S3, EC2, EMR, and Redshift
Build and maintain ETL/ELT data pipelines using PySpark for data ingestion and transformation
Develop and optimize big data processing jobs using PySpark on AWS EMR
Implement and manage data warehousing solutions focusing on Hive and modern table formats
Set up secure cloud infrastructure components including VPCs and security groups
Design and manage containerized applications on Amazon EKS
Optimize AWS resources for performance, cost, and efficiency
Implement best practices for data security and governance within AWS
Benefits
Opportunity to work with cutting-edge technologies in the AWS ecosystem
Collaborative environment with data scientists and engineers
Involvement in high-impact data solutions and infrastructure
Potential for professional development and certification support
Dynamic work location in Irving, Texas
Full Job Description
Role description
Job Title: AWS Pyspark Developer
Work Location : Irving,Texas
Job Summary
We are seeking a highly skilled and motivated AWS Certified Engineer to design build and optimize scalable data solutions within the Amazon Web Services AWS ecosystem The ideal candidate will have strong expertise in big data processing using PySpark and a deep understanding of data warehousing concepts including Hive and modern table formats like Iceberg This role involves developing deploying and managing robust efficient and secure data pipelines and analytics solutions on AWS leveraging core networking and compute services
Responsibilities
AWS Solution Design Implementation Design develop and deploy scalable and costeffective data solutions on AWS leveraging services such as S3 for data lakes EC2 EMR Glue Athena Lambda Redshift and Kinesis
Data Pipeline Development Build and maintain robust ETLELT data pipelines using PySpark for data ingestion transformation and loading into various data stores including those utilizing open table formats like Iceberg
Big Data Processing Develop and optimize big data processing jobs using PySpark on AWS EMR or AWS Glue handling large datasets efficiently and integrating with Iceberg table formats
Data Warehousing Design implement and manage data warehousing solutions including schema design data modeling and query optimization with a focus on Hive and modern data lake table formats like Iceberg for historical data and analytical queries
Cloud Infrastructure Networking Implement secure and robust cloud infrastructure components including VPCs subnets routing and security groups to ensure proper connectivity and isolation for data solutions
Containerized Workloads Design deploy and manage containerized data processing applications on Amazon Elastic Kubernetes Service EKS
Performance Tuning Optimization Optimize AWS resources and big data applications Spark Hive Iceberg for performance cost and efficiency
Data Governance Security Implement best practices for data security access control and compliance within AWS including IAM policies S3 bucket policies and encryption
Monitoring Troubleshooting Set up monitoring ing and logging for data pipelines and AWS infrastructure troubleshoot and resolve issues promptly
Automation Develop and maintain automation scripts using Python and shell scripting for infrastructure provisioning deployment and operational tasks
Collaboration Work closely with data scientists analysts and other engineering teams to understand data requirements and deliver reliable data solutions
Required Skills , Qualifications
AWS Certification Hold at least one AWS certification eg AWS Certified Solutions Architect Associate AWS Certified Data Analytics Specialty AWS Certified Developer Associate
AWS Services Expertise Handson experience with key AWS services for data processing and storage including
Storage S3 for data lakes EC2
Data Processing EMR Glue Athena Lambda
Networking VPC Subnets Routing Security Groups
Containerization EKS
Big Data Processing Strong proficiency in PySpark for developing complex data transformations and analytics
Data Lake Table Formats Practical experience with Apache Iceberg for managing and querying data lakes
Data Warehousing Indepth knowledge and practical experience with Apache Hive for data storage querying and schema management
Programming Languages
Python Expertlevel proficiency in Python for scripting data manipulation and AWS automation Boto3
Shell Scripting Proficient in shell scripting for automation and operational tasks
Database SQL Strong SQL skills for data querying and manipulation
Data Concepts Solid understanding of ETLELT processes data modeling distributed computing and data governance
Good to Have Skills
Containerization Orchestration Experience with Kubernetes for deploying and managing containerized applications
CICD Experience with CICD tools and practices eg AWS CodePipeline GitHub Actions GitLab CI for automating deployment of data solutions
Orchestration Experience with workflow orchestration tools like Apache Airflow
Version Control Proficient in using Git for source code management
Other Big Data Technologies Exposure to other big data technologies like Apache Kafka Flink or Presto