Position Summary:
The Clinical Data Engineer II/III owns clinical Data Operations: the data pipelines, data warehouse, and reporting infrastructure that transform clinical data, laboratory data, and sequencing results into validated, audit-ready clinical and business data. This position operates in a regulated environment in which accuracy and traceability are as important as well-architected code. The role combines hands-on pipeline and data-modeling work with ongoing production data support for our clinical study operations and is central to the integrity of the data that informs our research and clinical decisions.
Key Responsibilities:- Build and maintain ELT pipelines from clinical source systems - and non-standard sources such as ad hoc spreadsheets - into our data warehouse, and design ingestion frameworks that adapt to evolving report formats
- Support bulk data transfer initiatives to external partners and client-facing destinations, including the definition of the data contracts those partners depend on
- Develop, test, and refactor data transformation models across core clinical schemas, implementing precise, well-specified business logic, and integrate clinical data with other corporate data
- Define methods and metrics for assessing data quality, and monitor and maintain data quality across clinical datasets
- Own the reliability of clinical result data end to end, tracing issues to their root cause and strengthening controls through documented change management
- Strengthen the resilience of data contracts and design reconciliation logic that unifies data from multiple sources into a single, authoritative record
- Serve as a trusted partner to clinical and laboratory operations and external stakeholders, resolving data questions and improving accuracy at the source
- Build and maintain dashboards and monitoring, including those subject to formal system validation, and support the data needs of adjacent business systems
- Partner with the Regulatory/QA team to ensure clinical data flow and the data warehouse remain compliant with procedural and regulatory requirements
- Serve as a subject-matter expert on warehouse clinical data, providing technical guidance and training to research and development teams
- Provide project status updates to cross-functional stakeholders
Qualifications:
- 5+ years of hands-on data engineering experience with production data-transformation frameworks (dbt preferred) and a cloud data warehouse
- Software-engineering fundamentals and an expert understanding of database and data-warehouse concepts
- Deep SQL expertise, including the ability to identify subtle logic errors in complex queries
- Experience with ELT and data-integration tools
- Experience with designing, building and using containers
- Proficiency in Python or a comparable scripting language
- Familiarity with cloud identity and access-management concepts (key-pair authentication, OAuth/OIDC, workload identity federation)
- Proven ability to work in a regulated environment, including precise documentation, change-control processes, and audit trails
- Enjoys a role that balances building new data systems with keeping production systems reliable and trusted
- Strong written and verbal communication skills, with a track record of documenting your work
- Working knowledge of clinical or laboratory data systems is a plus
- Familiarity with clinical data standards (CDASH/CDISC) preferred
- Working knowledge of Good Clinical Practice (GCP) and the clinical trial process preferred
- Familiarity with cancer-related medical classifications and ontologies (e.g., ICD-10) a plus
On-site is preferred for this position, in either our San Diego or San Mateo office. The anticipated salary range for this position is $150,000 - $190,000. The actual salary will be dependent on various factors that may include experience level, knowledge, skills, and abilities.
ClearNote Health offers a competitive salary and benefits package.