Cherokee Nation Businesses

Healthcare Data Scientist

Cherokee Nation Businesses$110K — $130K *
US-AnywhereRemote in Washington, DC
Healthcare
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of hands-on experience with SQL for data transformation and analysis.
  • Proficiency in Python, particularly with pandas, for data processing.
  • Experience with PySpark for distributed data processing.
  • Hands-on experience in cloud-based data platforms, especially Azure Synapse.
  • Knowledge of ETL/ELT processes and data pipeline architectures.
  • Experience with structured and semi-structured data formats like CSV and JSON.
  • Familiarity with Git workflows for collaborative development.

Responsibilities

  • Develop and optimize data pipelines using SQL, Python, and PySpark.
  • Execute notebook workflows in Azure Synapse, including debugging and documentation.
  • Transform structured and semi-structured data, ensuring data quality.
  • Ingest and process healthcare data while preserving source lineage.
  • Conduct structured data validation and enforce quality checks.
  • Perform exploratory data analysis to identify patterns and data quality issues.
  • Collaborate with stakeholders to translate data requirements into usable datasets.

Benefits

  • Opportunities for professional development and training.
  • Flexible work schedule with options for remote work.
  • Collaborative environment that promotes cross-discipline teamwork.
  • Possibility of working on impactful healthcare data initiatives.
Full Job Description
JOB DESCRIPTION

ATA is seeking a Data Scientist to support data pipeline development, validation, and analysis within a cloud-based Health IT data platform. This role is hands-on and delivery-focused, with an emphasis on building reliable, reproducible data workflows using SQL, Python, and PySpark in an Azure Synapse environment.

A core expectation of this role is the ability to work across the full data lifecycle, from ingestion through transformation to final dataset delivery, while maintaining data quality and traceability. The ideal candidate is comfortable debugging data issues end-to-end, understands how data structure and join logic impact outputs, and applies disciplined validation and documentation practices. This role also supports exploratory data analysis and the development of derived datasets to enable analytics and downstream use cases. The position will work extensively with healthcare data originating from EHR systems and interface feeds, including HL7 v2 and FHIR data, clinical terminology, and source-to-target data mappings.

Key Responsibilities:
Data Pipeline Development

 Develop, maintain, and optimize data pipelines using SQL, Python (pandas), and

PySpark.

Advanced Technology Applications

 Execute and manage notebook-based workflows within Azure Synapse, including

debugging and documentation.

 Process and transform structured and semi-structured data in formats such as CSV,

JSON/NDJSON, and Parquet.

 Work within ETL/ELT pipelines across raw, curated, and production data layers.

 Ingest, profile, map, and transform healthcare data from EHR systems and interface

feeds while preserving source lineage and clinical context

Data Quality and Validation

 Perform structured data validation, including row counts, null checks, duplicate

detection, schema validation, and allowed value enforcement.

 Identify and resolve data quality issues such as schema drift, inconsistencies, and

transformation errors across pipeline stages.

 Apply repeatable testing and validation practices, including reproducing issues, verifying

fixes, and ensuring data reliability prior to downstream use.

 Validate source-to-target mappings and reconcile records across source and destination

systems during data conversion and migration activities.

Data Analysis and Dataset Development

 Conduct exploratory data analysis to identify patterns, anomalies, and data quality

concerns.

 Develop derived datasets to support reporting, analytics, and downstream data use

cases.

 Collaborate with stakeholders to translate data requirements into usable datasets and

metrics.

Data Modeling and Structure Awareness

 Interpret and work with data schemas, including column definitions, data types, primary

and composite keys, and table relationships.

 Manage dataset grain and understand how join strategies (e.g., one-to-one vs. one-to

many) impact row counts and outputs.

Debugging and Troubleshooting

 Trace data issues from source ingestion through transformation logic to final outputs.

 Use logs and debugging approaches to diagnose and resolve pipeline issues.

 Document data transformations, assumptions, mappings, and validation results in a

clear and consistent manner.

 Collaborate with engineers, analysts, and stakeholders to ensure data usability,

integrity, and alignment with requirements.

Advanced Technology Applications

 Communicate data issues, findings, and workflow updates with technical team

members.

Minimum Qualifications:

 Hands-on experience writing SQL queries for data transformation and analysis.

 Experience using Python (e.g., pandas) for data processing.

 Hands-on experience using PySpark for distributed data processing.

 Experience working within cloud-based data platforms, preferably Azure Synapse or

similar.

Understanding ETL/ELT concepts and data pipeline architecture.

 Experience working with structured and semi-structured data formats (CSV, JSON,

Parquet).

 Familiarity with Git and collaborative development workflows.

 Strong problem-solving and debugging skills across data pipelines.

 Ability to validate and ensure data quality through structured checks and testing

practices.

 Strong written and verbal communication skills.

Preferred Qualifications:

 Experience working with healthcare interoperability standards and message formats,

particularly HL7 v2 and FHIR.

 Familiarity with clinical terminology and code systems such as SNOMED CT, ICD-10,

LOINC, RxNorm, or CPT.

 Experience supporting large-scale EHR data conversion or migration efforts, particularly

between RPMS and Oracle Health/Cerneror a comparable legacy-to-modern EHR

migration.

 Understanding of clinical data mapping, terminology normalization, provenance, and

validation across source and target systems

General personal traits we know will connect well with the team:

 Dependable, self-directed, and able to meet commitments.

 Comfortable learning new subject areas and solving unfamiliar problems.

 Enjoys collaborating across technical and subject-matter disciplines.

 Pragmatic and able to select the appropriate tool for the requirement.

About Cherokee Nation Businesses

Cherokee Nation Businesses is a diversified holding company that manages a range of businesses and investments in various sectors, including gaming, hospitality, aerospace, real estate, technology, healthcare, natural resources, and more. The company is owned by the Cherokee Nation, the largest Native American tribe in the United States. Cherokee Nation Businesses is committed to creating economic opportunities and improving the quality of life for Cherokee citizens and the surrounding communities.
Learn more about Cherokee Nation Businesses
Size
7,000 employees
Industry

Similar Jobs

More Jobs at Cherokee Nation Businesses

More Healthcare Jobs

Find similar Healthcare Data Scientist jobs: