The Role
We are hiring a Data & Analytics Engineer to own the data and intelligence layer that powers clinical reporting, operational analytics, and AI workloads. You will design and build the large-scale Spark and Databricks pipelines that process device and clinical data, own data quality and lifecycle management across the platform, and build the MLOps and feature pipelines behind our AI-assisted workflows. This role sits at the center of engineering, product, clinical operations, and our AI program.
What You Will Do
- Design, build, and maintain large-scale ETL/ELT pipelines on Spark and Databricks that process device transmission and EMR data.
- Develop and tune PySpark jobs and Delta Lake tables for performance, reliability, and cost at scale.
- Build dimensional models and a conformed semantic layer that drive consistent metrics across clinical and operational reporting.
- Model curated data into the Azure SQL serving layer and develop datasets and embedded dashboards for clinic-facing and internal analytics.
- Design and own data quality frameworks, including validation rules, monitoring, anomaly detection, and remediation across the pipeline.
- Own data lifecycle management across the platform, including ingestion, retention, archival, lineage, and disposal aligned to HIPAA and compliance requirements.
- Build and operate MLOps pipelines for model deployment, versioning, monitoring, and retraining, supporting Atlas AI and other models.
- Build feature and measurement pipelines for AI-assisted workflows, establishing ground truth and tracking accuracy, precision, recall, and quality over time.
- Monitor model and data drift in production, detecting distribution shifts and performance degradation and triggering retraining or remediation.
- Orchestrate, schedule, and monitor data workflows for reliability at scale.
- Translate requirements from product, clinical operations, and the AI team into reliable, well-documented data assets.
What You Will Bring
- Five or more years building data pipelines and analytics solutions in production.
- Deep hands-on experience with Apache Spark and Databricks, including PySpark, Delta Lake, notebook-based development, and workflow orchestration.
- Strong Python skills for data transformation, automation, and pipeline development.
- Strong SQL skills, including T-SQL, query optimization, and dimensional data modeling.
- Experience migrating existing T-SQL stored procedures and queries into equivalent Spark and PySpark pipelines.
- Experience designing semantic or metrics layers and delivering analytics and dashboards to end users.
- Experience with Azure data services such as Azure SQL, Data Factory, Synapse, and Data Lake.
- Experience designing data quality frameworks, including validation, monitoring, and remediation.
- Experience with data lifecycle management, including retention, archival, governance, and lineage.
- Experience with MLOps, including model deployment, versioning, monitoring, drift detection, and feature pipelines.
- Clear written communication and the ability to document data assets for technical and non-technical audiences.
Nice to Have
- Experience with healthcare data and working under HIPAA and PHI handling requirements.
- Familiarity with MLflow, model registries, feature stores, or model monitoring tools.
- Experience with streaming or near-real-time pipelines, such as Spark Structured Streaming.
- Familiarity with LLM evaluation, agentic workflows, prompt engineering, or LLM-assisted development.
- Experience with DBT or a similar transformation framework.
- Familiarity with .NET framework and Dapper.
- Relevant cloud or data certifications, such as Databricks, Azure, or AWS data credentials.