Athelas

Staff Software Engineer, Data Warehouse

Athelas$150K — $180K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years of software engineering experience, specifically with data platforms at scale.
  • Proficient in modern data technologies including CDC and streaming solutions.
  • Experienced with data lake formats and MPP query engines.
  • Strong SQL skills and understanding of schema design and query optimization.
  • Experienced in managing production data infrastructure and its related practices.

Responsibilities

  • Own the data warehouse platform's entire architecture and operations.
  • Design and implement CDC pipelines for efficient data movement.
  • Architect a data lake on object storage with open formats.
  • Scale and optimize the MPP query layer for efficient data serving.
  • Develop a robust transformation layer with dbt for consistent metrics.
  • Establish orchestration and CI/CD tooling for data workflows.
  • Collaborate with Security and Compliance for handling sensitive data.

Benefits

  • Flexible working hours and remote-friendly policies.
  • Opportunities for professional development and growth.
  • Collaborative work environment with cross-functional teams.
  • Access to the latest tools and technologies in data engineering.
Full Job Description
About the Role

We're hiring a Staff Software Engineer to own Commure's data warehouse platform end-to-end. You'll design, build, and operate every layer of the stack:
  • CDC pipelines
  • Streaming transport
  • Schema governance and data contracts
  • Query and serving layer
  • Analytics platform

The core stack today: Debezium for CDC, StarRocks as our MPP query and serving engine, and dbt for transformation and modeling. This is a hands-on IC role with broad scope. You'll make architectural calls, write the code that matters most, and set the patterns other teams build on.

What You'll Do
  • Own the data warehouse platform end-to-end: CDC pipelines, data lake, query layer, transformation layer, and the analytics-facing tooling that sits on top.
  • Design and operate CDC pipelines with Debezium (and Kafka, Redpanda, or an equivalent streaming backbone) that move data from operational databases into the warehouse with low latency and high fidelity.
  • Architect the data lake on object storage using an open table format (Iceberg, Delta Lake, or Hudi) with Parquet, enabling both batch and streaming workloads and clean separation of storage from compute.
  • Run and scale StarRocks (or adjacent MPP/lakehouse engines) as the query and serving layer - schema design, materialized views, ingestion patterns, tuning, and cost/performance trade-offs.
  • Build the transformation layer with dbt: modeling standards, tests, documentation, and a semantic layer that gives every team a single source of truth for metrics.
  • Stand up orchestration (Airflow, Dagster, or similar) and the CI/CD, observability, and data-quality tooling that make the platform trustworthy day-to-day.
  • Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability so the platform meets HIPAA and SOC 2 bar by default.
  • Set patterns and conventions: schema contracts, ingestion patterns, and self-serve tooling - that let product and analytics teams build on the platform without needing you in the loop for every decision.


What You Have
  • 6+ years of software engineering experience, with significant time building or operating data platforms at scale.
  • Experience across the modern data stack: CDC (Debezium or equivalent), streaming (Kafka, Redpanda), data lake formats (Iceberg, Delta, Hudi), an MPP or lakehouse query engine (StarRocks, ClickHouse, Trino, Snowflake, Databricks), and dbt.
  • Fluent in SQL, schema design, query optimization, and reasoning about cost and latency trade-offs on large datasets.
  • Experience running production data infrastructure (orchestration, observability, on-call, data quality, and incident response).
Preferred
  • Direct experience with Debezium, StarRocks, and dbt in production.
  • Experience building semantic layers (dbt Semantic Layer, Cube) or data catalogs / lineage (DataHub, OpenMetadata, Amundsen).
  • Experience with HIPAA-regulated data (PHI handling, de-identification, and access governance).
  • Experience powering AI/ML workloads: feature stores, training-set curation, embedding pipelines, or retrieval systems.
  • Experience across multiple clouds (AWS, GCP, Azure), infrastructure-as-code (Terraform, Pulumi) and Kubernetes controllers.

About Athelas

Industry
Founded
2016

Similar Jobs

More Jobs at Athelas

More Enterprise Technology Jobs

Find similar Staff Software Engineer, Data Warehouse jobs: