AWS Lakehouse Data Engineer

System One Holdings, LLC

• $120K — $145K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in a relevant field or equivalent experience.
  • Six years of relevant experience in data engineering.
  • Hands-on experience with AWS-native data lake or lakehouse architectures on Amazon S3.
  • Strong experience with Python and PySpark for ETL/ELT pipeline development.
  • Proficiency in Apache Iceberg with knowledge of ACID transactions and schema evolution.
  • Advanced SQL skills for analytical tasks and data visualization.
  • Experience with data governance and access control using AWS services.

Responsibilities

  • Build and operate data pipelines from various sources including APIs and relational databases.
  • Design, implement, and optimize ETL/ELT pipelines for analytics-ready datasets.
  • Improve pipeline reliability through automated testing and operational monitoring.
  • Implement a Delta Lakehouse-style data platform on AWS using native services.
  • Manage a scalable lakehouse on Amazon S3 with Apache Iceberg.
  • Establish standardized environments for development and production.
  • Automate AWS provisioning and develop CI/CD pipelines for data components.

Benefits

  • Medical, dental, and vision health coverage options.
  • Spending accounts and life insurance plans.
  • Voluntary benefits offerings available.
  • Participation in a 401(k) plan.
Full Job Description
Job Title: AWS Lakehouse Data Engineer
Work Model: Remote - offsite

Responsibilities

  • Build and operate data pipelines (batch and streaming) from APIs, relational databases, file drops, event streams, and external partners.
  • Design, implement, test, and optimize ETL/ELT pipelines using Python and PySpark to produce analytics-ready datasets for reporting, visualization, and machine learning.
  • Implement incremental processing, change data capture (CDC), data contracts, schema validation, and reusable transformation frameworks.
  • Improve pipeline reliability through automated testing, orchestration, monitoring, retries, and operational runbooks.
  • Design and implement a Delta Lakehouse-style data platform on AWS using native services to provide Databricks-like capabilities for data engineering, analysis, and data visualization.
  • Build and manage a scalable lakehouse on Amazon S3 using Apache Iceberg and open columnar formats such as Apache Parquet.
  • Implement SQL-like table reliability features including ACID transactions, schema evolution, snapshot isolation, and time travel using Apache Iceberg.
  • Enable fast, interactive queries of lakehouse data via AWS-native services like Amazon Athena, Amazon EMR, AWS Glue, and Amazon Redshift where appropriate.
  • Optimize performance and cost through partitioning, file sizing, caching, lifecycle policies, and separating compute from storage.
  • Establish standardized environments for development, testing, and production with consistent configuration and controlled promotion.
  • Implement data governance, access control, lineage, and quality measures utilizing AWS-native services including AWS Lake Formation, AWS Glue Data Catalog, IAM, KMS.
  • Create a metadata repository with cataloging, ownership, classification, tagging, and discoverability features.
  • Enable end-to-end data lineage for audit and regulatory compliance.
  • Apply policy-based access, least privilege, data classification, retention, encryption, and secure handling controls.
  • Build data quality checks for freshness, completeness, validity, and anomaly detection, and publish SLA/SLO metrics.
  • Automate AWS provisioning with Infrastructure as Code (IaC), develop CI/CD pipelines for data components, and ensure platform observability.
  • Work collaboratively with cross-functional teams and maintain high-quality engineering documentation.


Requirements

  • Bachelor's degree in Engineering, Information Technology, Computer Science, Data Engineering, or related field, or four (4) years of equivalent practical experience.
  • Six (6) years of relevant experience.
  • Hands-on experience building AWS-native data lake or lakehouse architectures on Amazon S3.
  • Strong experience developing production ETL/ELT pipelines with Python and PySpark, including data modeling, transformation, and performance tuning.
  • Hands-on experience with Apache Iceberg, including ACID transactions, schema evolution, time travel, and query optimization.
  • Advanced SQL skills supporting analytical workloads, reporting, and data visualization.
  • Proven experience with data governance, cataloging, lineage, and access control using AWS services.
  • Knowledge of AWS security fundamentals: IAM, KMS, secrets management, network security, logging, SDLC.
  • Proven experience with Infrastructure as Code (IaC) and operating data platforms across environments.
  • Experience with CI/CD pipelines for data workflows with testing, deployment, environment promotion, and rollback.
  • Troubleshooting distributed data workloads, performance optimization, and cost management skills.
  • Excellent collaboration and communication skills to coordinate with cross-team stakeholders.


Would Be Nice to Have

  • Experience with Databricks, Delta Lake, migrating workloads to AWS-native services, and Apache Iceberg.
  • Familiarity with AWS Step Functions, MWAA, Kinesis, DMS, Lambda, MSK, or similar services.
  • Experience with modern DevOps tools: Git, Terraform, CloudFormation, Jenkins, CodePipeline, GitHub Actions, Docker.
  • Knowledge of BI and visualization tools like Amazon QuickSight, Tableau, Power BI.
  • Familiarity with AI-assisted coding tools such as GitHub Copilot, ChatGPT, Cursor, or Kiro.
  • Knowledge of graph modeling, ontology, taxonomy, entity resolution, and hybrid retrieval techniques.


System One not only serves as a valued partner for our clients, but we offer eligible employees health and welfare benefits coverage options including medical, dental, vision, spending accounts, life insurance, voluntary plans, as well as participation in a 401(k) plan.

#LI-KA1
#M1

Ref: #851-Rockville-S1

Similar Jobs

More Jobs at System One Holdings, LLC

  • AWS Lakehouse Data Engineer
    $120K — $145K *
    Mclean, VA 22101 (Fairfax County)
    Information Technology
    In-Person
  • Project Manager
    $90K — $110K *
    Beaverton, OR 97007 (Washington County)
    Telecommunications & Hardware
    In-Person
  • Estimator
    $120K — $140K *
    Barnegat, NJ 08005 (Ocean County)
    Real Estate & Construction
    In-Person
  • Senior Contact Center Engineer
    $95K — $115K *
    Lafayette, LA 70506 (Lafayette County)
    Telecommunications & Hardware
    In-Person
  • Senior Cloud Administrator (onsite)
    $110K — $130K *
    Washington, DC 20011 (District Of Columbia County)
    Enterprise Technology
    In-Person

More Information Technology Jobs

Find similar AWS Lakehouse Data Engineer jobs: