EPAM Systems

Senior Data Software Engineer with AWS and Terraform

EPAM Systems$130K — $155K *
US-AnywhereRemote in Georgia, US
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of production-level experience in Python for data engineering
  • Expertise in PySpark and DataFrame API for distributed data processing
  • Advanced knowledge of Snowflake and data warehousing techniques
  • Proficient in AWS Glue job configuration and data quality rules
  • Experience with AWS services such as S3, IAM, Lambda, and Athena
  • Strong understanding of data quality frameworks and reconciliation processes
  • English proficiency at B2 level or higher

Responsibilities

  • Design and optimize scalable ETL/ELT pipelines using PySpark on AWS Glue
  • Configure Glue components including jobs, crawlers, and the Glue Data Catalog
  • Tune performance and cost of Glue jobs through partitioning and shuffle behavior
  • Model and load curated datasets into Snowflake with multiple layers
  • Implement automated data quality frameworks and conduct validation checks
  • Conduct row-count reconciliation and handle anomaly detection
  • Engage directly with clients for requirements and design discussions

Benefits

  • High autonomy within a senior consultant role
  • Opportunity to work directly with client stakeholders
  • Access to modern cloud technologies and platforms
  • Participation in CI/CD and code review processes
  • Focus on innovative data solutions and quality engineering
Full Job Description
We are seeking a Senior Data Software Engineer to join a client-facing delivery team building and hardening cloud-native data pipelines on AWS as part of a data platform modernization program. The role involves ingesting and transforming large datasets with PySpark on AWS Glue and delivering curated, validated data into Snowflake, with a core focus on data quality, validation, and reconciliation for downstream analytics. This position is delivered at a Senior Consultant level with high autonomy and direct client stakeholder communication. Responsibilities Design, build, and optimize scalable batch and incremental ETL/ELT pipelines using PySpark on AWS Glue Configure Glue jobs, crawlers, triggers, connections, bookmarks, workflows, and the Glue Data Catalog Tune workers, partitioning, and shuffle behavior for cost and performance optimization Model and load curated datasets into Snowflake with staging, transformation, and publishing layers Implement automated data quality and validation frameworks, including schema/contract enforcement and null/uniqueness/referential checks Develop row-count and financial reconciliation processes, anomaly detection, and quarantine/reject handling Configure and extend Glue Data Quality (DQDL) rules per requirements Write clean, modular, testable Python with unit/integration tests and reusable libraries Integrate pipelines with AWS services such as S3, IAM, Lambda, Athena, CloudWatch, Step Functions, and Secrets Manager Instrument observability through logging, metrics, alerting, and pipeline SLA monitoring Participate in code reviews, CI/CD automation, and documentation Engage directly with client stakeholders in requirements refinement, design walkthroughs, status reporting, and act as technical advisor within the workstream Requirements 3+ years of experience with Python for production-level data engineering, including OOP and functional patterns Expertise in PySpark for distributed data processing and the DataFrame API Advanced proficiency in Snowflake, including data warehousing and staging/transformation layers Skills in AWS Glue, including job configuration, crawlers, Data Catalog, and DQDL Background in data quality engineering, including validation frameworks and reconciliation Proficiency in AWS services including S3, IAM, Lambda, Athena, and CloudWatch English proficiency at B2 level or higher Nice to have Familiarity with Generative AI / LLM concepts Knowledge of Airflow / Step Functions orchestration Familiarity with Great Expectations or similar data quality frameworks Knowledge of Terraform / CloudFormation

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Information Technology Jobs

Find similar Senior Data Software Engineer with AWS and Terraform jobs: