Senior Data Engineer (On Site, Washington, DC)

Agile5 Technologies, Inc.

$120K — $145K *
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-10 years of relevant experience in data engineering, with specific focus on ETL processes and data migration.
  • Strong proficiency in Python, PySpark, and SQL for developing scalable data pipelines.
  • Experience with Databricks, AWS services (particularly in GovCloud environments), and transitioning data from legacy systems to cloud-native architectures.
  • Deep understanding of data governance principles, particularly unity catalog and role-based access control (RBAC).
  • Educational background in Computer Science, Information Technology, or a related discipline.

Responsibilities

  • Lead technical aspects of data migration operations, ensuring accuracy and reliability.
  • Migrate existing data architectures (e.g., Hive) to Delta Lake format on AWS S3.
  • Convert complex ETL mappings from Informatica to Python/PySpark while maintaining data integrity.
  • Design and implement change data capture (CDC) pipelines for real-time data synchronization.
  • Manage and maintain Databricks workspaces and configure CI/CD pipelines for Python deployments.
  • Ensure data governance compliance by implementing Unity Catalog functionalities.
  • Mentor junior engineers and engage in Agile practices for project delivery.

Benefits

  • Opportunity to work in a secure environment focused on critical data systems.
  • Mentorship opportunities to guide and develop mid-level engineering personnel.
  • Engagement in complex, strategic projects within a recognized enterprise data modernization effort.
  • Access to ongoing professional development through certifications and technical training.
Full Job Description
Description: The Senior Data Engineer serves as the technical lead for complex data migration and pipeline engineering initiatives within enterprise data lakehouse modernization efforts. This role leads the conversion of legacy ETL artifacts into scalable Python/PySpark code, migrates database systems to cloud-native Delta Lake architectures, and implements robust data governance and automated reconciliation frameworks. Operating in a secure environment, the ideal candidate will manage Databricks infrastructure, design CI/CD pipelines, and mentor engineering personnel while ensuring complete data accuracy.

Senior Data Engineer Job Duties:
  • Serve as the primary technical lead executing daily migration operations, including data profiling, schema analysis, pipeline conversion, automated reconciliation, and production deployment validation.
  • Migrate legacy Hive tables to Delta Lake format on AWS S3 storage using Databricks ingestion tools.
  • Convert High and Medium complexity Informatica mappings to Python/PySpark code while preserving all business logic, data quality checks, and data cleansing routines.
  • Design and implement Change Data Capture (CDC) pipelines to maintain data synchronization between legacy and target environments during migration.
  • Configure and maintain Databricks workspaces, clusters, notebooks, and jobs within AWS GovCloud environments.
  • Implement Unity Catalog data governance, including RBAC, column-level encryption, audit trails, and data lineage tracking.
  • Build and execute automated data reconciliation scripts validating 100% migration accuracy.
  • Develop and maintain CI/CD pipelines for Python code deployment using Azure DevOps, checking in all converted code with comprehensive documentation.
  • Integrate the Databricks lakehouse with downstream applications (such as Power BI and ESRI) and configure SSO integrations.
  • Access legacy enclave environments for data profiling and side-by-side validation while collaborating daily with database managers and analysts.
  • Mentor mid-level engineering staff on platform operations and migration methodologies, participate in Agile ceremonies, and support system acceptance demonstrations.
  • Performs other duties as assigned.

Security Clearance Requirements:
  • Public Trust / Tier 4 Eligible: No clearance required to apply; must be a U.S. citizen willing to undergo a background check to obtain a Public Trust / Tier 4 clearance.

Experience Requirements:
  • Minimum experience required varies by degree level: PhD with 4 years; Master's degree with 8 years; Bachelor's degree with 10 years; or High School Diploma with 14 years of relevant experience.
  • Hands-on experience with Databricks (workspace administration, notebook development, job scheduling) and proficiency in Python, PySpark, and SQL for large-scale pipeline development.
  • Experience migrating data from legacy platforms (Hive, Hadoop, Oracle) to cloud-native platforms, utilizing Delta Lake or Apache Iceberg table formats.
  • Practical experience with AWS services (S3, IAM, RDS, GovCloud), CI/CD pipeline tools (Azure DevOps, Jenkins, GitHub Actions), and enterprise ETL/ELT frameworks.

Education Requirements: Bachelor's degree in Computer Science, Data Engineering, Information Technology, or a related technical field (or equivalent combination of education and experience).

Desired Skills / Qualifications:
  • Databricks Certified Data Engineer (Associate or Professional).
  • Direct experience converting Informatica PowerCenter mappings to Python/PySpark code.
  • Experience with Unity Catalog data governance, Cloudera Hive/HiveQL, Power BI/ESRI integrations, and CDC methodologies.
  • Familiarity with FISMA High or FedRAMP High security requirements and SAML-based SSO/Okta integrations.

Location: Washington, DC

Status: Full time

Schedule: Day shift, Monday-Friday

Physical Requirements: Must be able to remain in a stationary position for long durations of time. Also, must be able to continuously operate a computer and other office productivity machinery.

Travel Required: No

This job description is subject to change at any time.

Similar Jobs

More Jobs at Agile5 Technologies, Inc.

More Enterprise Technology Jobs

Find similar Senior Data Engineer (On Site, Washington, DC) jobs: