Data, ML Ops and Analytics Engineer

Octagos Health

$90K — $130K *
Healthcare
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience building data pipelines and analytics solutions in production.
  • Hands-on expertise with Apache Spark and Databricks, including PySpark and Delta Lake.
  • Strong Python skills for data transformation and automation.
  • Proficient in SQL, including T-SQL and query optimization.
  • Experience designing semantic layers and delivering relevant analytics dashboards.
  • Knowledge of Azure data services like Azure SQL and Data Factory.
  • Demonstrated experience with data quality frameworks and lifecycle management.

Responsibilities

  • Design, build, and maintain large-scale ETL/ELT pipelines on Spark and Databricks.
  • Develop and optimize PySpark jobs and Delta Lake tables.
  • Build dimensional models that establish consistency in metrics reporting.
  • Model data into Azure SQL for analytics and dashboards.
  • Design data quality frameworks encompassing validation and monitoring processes.
  • Oversee data lifecycle management aligned with compliance requirements.
  • Build and operate MLOps pipelines for AI model deployment and monitoring.

Benefits

  • Opportunity to work with cutting-edge AI technologies in a healthcare context.
  • Collaboration across multiple disciplines including engineering, product, and clinical operations.
  • Engagement with large-scale datasets that have real-world impact on patient care.
  • Strong emphasis on data quality and compliance, particularly under HIPAA regulations.
Full Job Description
The Role

We are hiring a Data & Analytics Engineer to own the data and intelligence layer that powers clinical reporting, operational analytics, and AI workloads. You will design and build the large-scale Spark and Databricks pipelines that process device and clinical data, own data quality and lifecycle management across the platform, and build the MLOps and feature pipelines behind our AI-assisted workflows. This role sits at the center of engineering, product, clinical operations, and our AI program.

What You Will Do
  • Design, build, and maintain large-scale ETL/ELT pipelines on Spark and Databricks that process device transmission and EMR data.
  • Develop and tune PySpark jobs and Delta Lake tables for performance, reliability, and cost at scale.
  • Build dimensional models and a conformed semantic layer that drive consistent metrics across clinical and operational reporting.
  • Model curated data into the Azure SQL serving layer and develop datasets and embedded dashboards for clinic-facing and internal analytics.
  • Design and own data quality frameworks, including validation rules, monitoring, anomaly detection, and remediation across the pipeline.
  • Own data lifecycle management across the platform, including ingestion, retention, archival, lineage, and disposal aligned to HIPAA and compliance requirements.
  • Build and operate MLOps pipelines for model deployment, versioning, monitoring, and retraining, supporting Atlas AI and other models.
  • Build feature and measurement pipelines for AI-assisted workflows, establishing ground truth and tracking accuracy, precision, recall, and quality over time.
  • Monitor model and data drift in production, detecting distribution shifts and performance degradation and triggering retraining or remediation.
  • Orchestrate, schedule, and monitor data workflows for reliability at scale.
  • Translate requirements from product, clinical operations, and the AI team into reliable, well-documented data assets.

What You Will Bring
  • Five or more years building data pipelines and analytics solutions in production.
  • Deep hands-on experience with Apache Spark and Databricks, including PySpark, Delta Lake, notebook-based development, and workflow orchestration.
  • Strong Python skills for data transformation, automation, and pipeline development.
  • Strong SQL skills, including T-SQL, query optimization, and dimensional data modeling.
  • Experience migrating existing T-SQL stored procedures and queries into equivalent Spark and PySpark pipelines.
  • Experience designing semantic or metrics layers and delivering analytics and dashboards to end users.
  • Experience with Azure data services such as Azure SQL, Data Factory, Synapse, and Data Lake.
  • Experience designing data quality frameworks, including validation, monitoring, and remediation.
  • Experience with data lifecycle management, including retention, archival, governance, and lineage.
  • Experience with MLOps, including model deployment, versioning, monitoring, drift detection, and feature pipelines.
  • Clear written communication and the ability to document data assets for technical and non-technical audiences.

Nice to Have
  • Experience with healthcare data and working under HIPAA and PHI handling requirements.
  • Familiarity with MLflow, model registries, feature stores, or model monitoring tools.
  • Experience with streaming or near-real-time pipelines, such as Spark Structured Streaming.
  • Familiarity with LLM evaluation, agentic workflows, prompt engineering, or LLM-assisted development.
  • Experience with DBT or a similar transformation framework.
  • Familiarity with .NET framework and Dapper.
  • Relevant cloud or data certifications, such as Databricks, Azure, or AWS data credentials.

Similar Jobs

More Jobs at Octagos Health

More Healthcare Jobs

Find similar Data, ML Ops and Analytics Engineer jobs: