Position Overview:
*This position is contingent upon contract award*
The Lead Data Engineer is responsible for PDE processing pipelines, data quality, and IDR integration for the Drug Data Processing System (DDPS) and Payment Reconciliation System (PRS) O&M contract. This individual owns the end-to-end PDE data pipeline - receiving, validating, editing, and storing approximately 9 to 10 million Prescription Drug Event records daily - and maintains all vendor reference data integrations including FDA, NCPDP, FDB, MediSpan, NPPES, and OIG data sources.
This is a remote position; however, Cognitive hires only in the following designated U.S. states based on contract and business requirements: VA, DC, MD, TN, FL, AZ, CO, OR, and TX.
Key Responsibilities:
- Own summary records of prescription drug transactions, also known as pharmacy drug events (PDE), ingestion, validation, editing, and storage pipelines.
- Apply and maintain CMS-defined business rule validations, returning edit results to plan sponsors, and minimizing PDEs requiring manual analysis.
- Engineer and maintain all vendor reference data edits from a variety of sources for drug event record, pharmacy data and files, exclusion and preclusion list; provide health plan and CMS status updates on PDE submissions.
- Build, maintain, and optimize Apache Airflow and PySpark transformations across variety of technologies, including S3, Redshift Serverless, Snowflake, and Databricks.
- Maintain system compatibility with other CMS data model changes.
- Monitor PDE submissions for adjustment trends and informational edit rates; provide analytical capability to link rejected PDEs with IDR records; provide beneficiary and plan-level cost aggregations to PRS for year-end reconciliation.
- Support data transfer projects and new data share integrations with CMS and downstream stakeholders including API and multi-cloud approaches as directed by the CMS.
- Maintain target unit test coverage for all new pipeline code; support BDD and TDD practices; track and surface real-time PDE processing metrics including volumes, error rates, and reconciliation accuracy.
- Operate within the CMS Lean-Agile Release Train adhering to the Iteration Schedule; deliver code through GitHub, Jenkins, JFrog Artifactory/XRay, SonarQube, and Snyk CI/CD pipeline.
Qualifications:
- Bachelor's degree in Computer Science, Data Engineering, or related field; 8 or more years in data engineering, including 4 or more years building and operating high-volume batch ETL systems.
- Expert-level Python andPySpark
- Production experience with Apache Airflow (AWS MWAA) for pipeline orchestration at scale.
- Production experience with AWS Redshift Serverless, Snowflake, and Amazon Athena for data warehousing and query at scale.
- Hands-on experience with Databricks (Notebooks, Jobs) for large-scale data processing.
- Strong SQL skills for multi-source data transformation, editing logic, and reconciliation of data pipelines.
- Experience with AWS S3, SQS, Lambda, and Event Bridge for data ingestion and event-driven orchestration.
- Experience with Linux environments including RHEL, CentOS, and Amazon Linux 2, and shell scripting in Bash.
- Demonstrated experience operating data pipelines under strict data quality, accuracy, and timeliness SLAs in a federal or healthcare environment.
- Experience with version control and CI/CD integration using GitHub, Jenkins, and JFrog Artifactory.
- Ability to pass CMS and internal required background checks for public trust.