Lead Data Engineer - Remote (WFH)

Cognitive Medical Systems

$120K — $145K *
US-AnywhereRemote in United States
Healthcare
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's degree in Computer Science, Data Engineering, or related field; 8+ years in data engineering, including 4+ years in high-volume batch ETL systems.
  • Expert-level proficiency in Python and PySpark.
  • Extensive production experience with Apache Airflow (AWS MWAA).
  • Hands-on experience with AWS Redshift Serverless, Snowflake, and Amazon Athena.
  • Familiarity with Databricks for large-scale data processing.
  • Strong SQL skills for multi-source data transformations and reconciliation.
  • Experience with AWS services such as S3, SQS, Lambda, and Event Bridge.
  • Experience working in Linux environments and shell scripting in Bash.
  • Proven track record of meeting strict data quality and timeliness SLAs in federal or healthcare environments.

Responsibilities

  • Own the ingestion, validation, editing, and storage of Prescription Drug Event (PDE) records.
  • Apply CMS-defined business rule validations and minimize manual analysis of PDEs.
  • Engineer and maintain vendor reference data edits from multiple sources.
  • Build and maintain Apache Airflow and PySpark transformations across various technologies.
  • Monitor PDE submissions and analyze trends to link rejected PDEs with IDR records.
  • Support data transfer projects and new integrations with CMS and other stakeholders.
  • Maintain target unit test coverage and support testing practices for new pipeline code.
  • Deliver code through CI/CD pipeline tools and operate within a Lean-Agile Release Train.

Benefits

  • Remote work opportunity with designated states for hiring.
  • Engagement with high-impact data-related projects in the healthcare sector.
  • Opportunity to work with cutting-edge technologies like AWS, Databricks, and Apache Airflow.
  • Be part of an agile team focused on delivering quality-driven data solutions.
  • Potential for career advancement in a growing technology-driven firm.
Full Job Description
Position Overview:

*This position is contingent upon contract award*

The Lead Data Engineer is responsible for PDE processing pipelines, data quality, and IDR integration for the Drug Data Processing System (DDPS) and Payment Reconciliation System (PRS) O&M contract. This individual owns the end-to-end PDE data pipeline - receiving, validating, editing, and storing approximately 9 to 10 million Prescription Drug Event records daily - and maintains all vendor reference data integrations including FDA, NCPDP, FDB, MediSpan, NPPES, and OIG data sources.

This is a remote position; however, Cognitive hires only in the following designated U.S. states based on contract and business requirements: VA, DC, MD, TN, FL, AZ, CO, OR, and TX.

Key Responsibilities:
  • Own summary records of prescription drug transactions, also known as pharmacy drug events (PDE), ingestion, validation, editing, and storage pipelines.
  • Apply and maintain CMS-defined business rule validations, returning edit results to plan sponsors, and minimizing PDEs requiring manual analysis.
  • Engineer and maintain all vendor reference data edits from a variety of sources for drug event record, pharmacy data and files, exclusion and preclusion list; provide health plan and CMS status updates on PDE submissions.
  • Build, maintain, and optimize Apache Airflow and PySpark transformations across variety of technologies, including S3, Redshift Serverless, Snowflake, and Databricks.
  • Maintain system compatibility with other CMS data model changes.
  • Monitor PDE submissions for adjustment trends and informational edit rates; provide analytical capability to link rejected PDEs with IDR records; provide beneficiary and plan-level cost aggregations to PRS for year-end reconciliation.
  • Support data transfer projects and new data share integrations with CMS and downstream stakeholders including API and multi-cloud approaches as directed by the CMS.
  • Maintain target unit test coverage for all new pipeline code; support BDD and TDD practices; track and surface real-time PDE processing metrics including volumes, error rates, and reconciliation accuracy.
  • Operate within the CMS Lean-Agile Release Train adhering to the Iteration Schedule; deliver code through GitHub, Jenkins, JFrog Artifactory/XRay, SonarQube, and Snyk CI/CD pipeline.

Qualifications:
  • Bachelor's degree in Computer Science, Data Engineering, or related field; 8 or more years in data engineering, including 4 or more years building and operating high-volume batch ETL systems.
  • Expert-level Python andPySpark
  • Production experience with Apache Airflow (AWS MWAA) for pipeline orchestration at scale.
  • Production experience with AWS Redshift Serverless, Snowflake, and Amazon Athena for data warehousing and query at scale.
  • Hands-on experience with Databricks (Notebooks, Jobs) for large-scale data processing.
  • Strong SQL skills for multi-source data transformation, editing logic, and reconciliation of data pipelines.
  • Experience with AWS S3, SQS, Lambda, and Event Bridge for data ingestion and event-driven orchestration.
  • Experience with Linux environments including RHEL, CentOS, and Amazon Linux 2, and shell scripting in Bash.
  • Demonstrated experience operating data pipelines under strict data quality, accuracy, and timeliness SLAs in a federal or healthcare environment.
  • Experience with version control and CI/CD integration using GitHub, Jenkins, and JFrog Artifactory.
  • Ability to pass CMS and internal required background checks for public trust.

Similar Jobs

More Jobs at Cognitive Medical Systems

More Healthcare Jobs

Find similar Lead Data Engineer - Remote (WFH) jobs: