Data Engineer - Legacy Systems & AI Workflows

Groundswell Agriculture Festival

• $89K — $175K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Computer Engineering, Mathematics, Statistics, or related field.
  • 5+ years of professional data engineering experience, with a focus on building production pipelines.
  • Proficiency in programming languages such as Python, Java, or R for data engineering tasks.
  • Experience with SOAP APIs and legacy systems for data interactions.
  • Background in implementing data-quality frameworks and monitoring for production pipelines.
  • Knowledge of data security practices and compliance in regulated environments.
  • Experience with structured data formats like JSON, XML, CSV, and Parquet.

Responsibilities

  • Design, develop, and operate ETL/ELT pipelines for data ingestion from various sources.
  • Onboard legacy data sources by reviewing and validating schemas and operational procedures.
  • Create automated data-quality checks to monitor data completeness and accuracy.
  • Automate deployments and pipeline operations using APIs and workflow management.
  • Troubleshoot failures in distributed data systems and communicate effectively about issues.
  • Collaborate with engineers and stakeholders to translate mission needs into actionable data products.

Benefits

  • Comprehensive medical, dental, and vision plans.
  • Flexible Spending Account options available.
  • 4% 401K Match with immediate vesting.
  • Generous Paid Time Off policy.
  • Tuition reimbursement and professional development opportunities.
  • Flexible work schedule offered.
  • On-site gym and childcare options available.
Full Job Description
What You'll do:

Groundswell is seeking an experienced Data Engineer to build secure, reliable data capabilities for a mission-focused platform that connects authoritative sources, operational data, and analytical services. This role spans data ingestion, integration, quality, governance, and pipeline operations. You will help make complex data usable and trustworthy by designing repeatable workflows that preserve provenance, enforce access controls, and support decision-ready applications and AI-enabled capabilities in controlled environments.

What You'll Do
  • Design, develop, and operate batch and streaming ETL/ELT pipelines that ingest data from multiple structured and semi-structured sources into secure cloud data platforms.
  • Onboard legacy data sources by reviewing schemas, mappings, interfaces, validation rules, ownership, and operational support procedures.
  • Build automated data-quality checks for completeness, accuracy, consistency, timeliness, and referential integrity, with actionable monitoring and alerting.
  • Automate deployments and pipeline operations using SOAP APIs, workflow orchestration, and environment-specific configurations.
  • Troubleshoot failures across distributed data systems using logs, metrics, lineage, and operational signals; communicate root causes and recovery plans clearly.
  • Collaborate with software engineers, cloud engineers, data scientists, security teams, architects, and customer stakeholders to translate mission needs into measurable data products.

Required Qualifications
  • Bachelor's degree in Computer Science, Computer Engineering, Mathematics, Statistics, or a related technical field.
  • 5+ years of professional data engineering experience, including production pipeline development and operations.
  • Proficiency in Python, Java, R, or other programming language used for data engineering.
  • Experience with legacy systems that utilize SOAP APIs for data interaction.
  • Experience implementing data-quality frameworks, validation processes, monitoring, and alerting for production pipelines.
  • Working knowledge of data security, access controls, encryption, auditability, and governance in a regulated, restricted, or compliance-oriented environment.
  • Experience working with APIs (SOAP/REST) and structured data formats such as JSON, XML, CSV, and Parquet.
  • Ability to document technical decisions and collaborate with multidisciplinary teams and customer stakeholders.
  • Must be a U.S. Citizen per contract requirements.
  • Must be able to obtain and maintain a Public Trust Clearance in accordance with contract requirements.


Preferred Qualifications
  • Experience building preprocessing, chunking, filtering, metadata-enrichment, or evaluation pipelines for LLM and other AI/ML workflows.
  • Experience operating data platforms with segmented networks, limited connectivity, strict change control, or formal authorization requirements.
  • Strong SQL expertise, including query optimization, data modeling, joins, window functions, and analysis of large datasets.
  • Experience with Jupyter Notebook or other equivalent tools for analyzing data.
  • Experience with Python libraries such as numpy, pandas, and other libraries like this.
  • Hands-on experience with at least one major cloud provider, with AWS experience strongly preferred.
  • Extra consideration given to those with C3.ai experience.
  • Active Public Trust Clearance.
  • Preference given to candidates local to the Washington, DC metro area


Certifications
  • AWS Data Engineer Certification, Databricks, cloud data engineering, or other relevant professional certification is preferred.


Skills:

Certification:

Why You'll Never Want to Leave:
  • Comprehensive medical, dental, and vision plans
  • Flexible Spending Account
  • 4% 401K Match (immediate vesting)
  • Paid Time Off
  • Tuition reimbursement, certification programs, and professional development
  • Flexible work schedule
  • On-site gym and childcare option


The salary range for this role takes into account the wide range of factors that are considered in making compensation decisions, including but not limited to skill sets, experience and training; licensure and certifications; and other business and organizational needs. The disclosed range estimate has not been adjusted for any applicable geographic differential associated with the location at which the position may be filled. At Groundswell, it is not typical for an individual to be hired at or near the top of the range for their role, and compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range is:
$89,886.00 - $175,444.00

NOTE: Groundswell does not accept unsolicited resumes through or from search firms or staffing agencies. All unsolicited resumes will be considered the property of Groundswell, and Groundswell will not be obligated to pay a placement fee.

Similar Jobs

  • Senior EDI Support Specialist
    $72K — $108K *
    UST
    Washington, DC 20011 (District Of Columbia County)
  • Mid Data Engineer
    $110K — $132K *
    Ignite Digital Services
    Washington, DC 20011 (District Of Columbia County)
  • Platform Engineer
    $110K — $130K *
    JLGOV
    Washington, DC 20011 (District Of Columbia County)
  • Data Engineer
    $110K — $130K *
    Black Canyon Consulting
    Bethesda, MD 20817 (Montgomery County)
  • ETL DEVELOPER (CI POLYGRAPH REQUIRED)
    $110K — $130K *
    NorthHill Technology
    Reston, VA 20191 (Fairfax County)
  • Data Engineer
    $110K — $130K *
    Absolute Business Solutions Corp.
    Chantilly, VA 20152 (Loudoun County)

More Jobs at Groundswell Agriculture Festival

More Information Technology Jobs

Find similar Data Engineer - Legacy Systems & AI Workflows jobs: