Athelas

Staff Software Engineer, Data Warehouse

Athelas$140K — $170K *
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 6+ years in software engineering, with a focus on data platform operations
  • Deep knowledge of modern data stack components including CDC, streaming, data lakes, and MPP engines
  • Proficient in SQL, schema design, and query optimization for large datasets
  • Experience managing production data infrastructure and associated tools
  • Preferred: Practical knowledge of Debezium, StarRocks, and dbt in real-world settings
  • Familiarity with HIPAA compliance and managing PHI/PII data
  • Exposure to AI/ML workloads and multi-cloud environments.

Responsibilities

  • Own the entire data warehouse platform from pipelines to analytics tools
  • Design and implement CDC pipelines using Debezium and streaming technologies
  • Architect and manage the data lake utilizing open table formats for flexibility
  • Scale StarRocks as the primary query layer, focusing on design and performance tuning
  • Develop the transformation layer using dbt for consistent metrics management
  • Establish orchestration and CI/CD processes alongside data quality monitoring
  • Collaborate on security and compliance measures for sensitive data handling
  • Define conventions and standards to empower teams to build independently.

Benefits

  • Flexible work environment with options for remote collaboration
  • Continuous learning and development opportunities
  • Strong focus on innovation and leveraging cutting-edge technologies
  • Robust support for compliance and data governance practices
  • Collaborative and inclusive company culture fostering team autonomy.
Full Job Description
About the Role

We're hiring a Staff Software Engineer to own Commure's data warehouse platform end-to-end. You'll design, build, and operate every layer of the stack:
  • CDC pipelines
  • Streaming transport
  • Schema governance and data contracts
  • Query and serving layer
  • Analytics platform

The core stack today: Debezium for CDC, StarRocks as our MPP query and serving engine, and dbt for transformation and modeling. This is a hands-on IC role with broad scope. You'll make architectural calls, write the code that matters most, and set the patterns other teams build on.

What You'll Do
  • Own the data warehouse platform end-to-end: CDC pipelines, data lake, query layer, transformation layer, and the analytics-facing tooling that sits on top.
  • Design and operate CDC pipelines with Debezium (and Kafka, Redpanda, or an equivalent streaming backbone) that move data from operational databases into the warehouse with low latency and high fidelity.
  • Architect the data lake on object storage using an open table format (Iceberg, Delta Lake, or Hudi) with Parquet, enabling both batch and streaming workloads and clean separation of storage from compute.
  • Run and scale StarRocks (or adjacent MPP/lakehouse engines) as the query and serving layer - schema design, materialized views, ingestion patterns, tuning, and cost/performance trade-offs.
  • Build the transformation layer with dbt: modeling standards, tests, documentation, and a semantic layer that gives every team a single source of truth for metrics.
  • Stand up orchestration (Airflow, Dagster, or similar) and the CI/CD, observability, and data-quality tooling that make the platform trustworthy day-to-day.
  • Partner with Security and Compliance on PHI/PII handling, access controls, lineage, and auditability so the platform meets HIPAA and SOC 2 bar by default.
  • Set patterns and conventions: schema contracts, ingestion patterns, and self-serve tooling - that let product and analytics teams build on the platform without needing you in the loop for every decision.


What You Have
  • 6+ years of software engineering experience, with significant time building or operating data platforms at scale.
  • Experience across the modern data stack: CDC (Debezium or equivalent), streaming (Kafka, Redpanda), data lake formats (Iceberg, Delta, Hudi), an MPP or lakehouse query engine (StarRocks, ClickHouse, Trino, Snowflake, Databricks), and dbt.
  • Fluent in SQL, schema design, query optimization, and reasoning about cost and latency trade-offs on large datasets.
  • Experience running production data infrastructure (orchestration, observability, on-call, data quality, and incident response).
Preferred
  • Direct experience with Debezium, StarRocks, and dbt in production.
  • Experience building semantic layers (dbt Semantic Layer, Cube) or data catalogs / lineage (DataHub, OpenMetadata, Amundsen).
  • Experience with HIPAA-regulated data (PHI handling, de-identification, and access governance).
  • Experience powering AI/ML workloads: feature stores, training-set curation, embedding pipelines, or retrieval systems.
  • Experience across multiple clouds (AWS, GCP, Azure), infrastructure-as-code (Terraform, Pulumi) and Kubernetes controllers.

About Athelas

Industry
Founded
2016

Similar Jobs

More Jobs at Athelas

More Enterprise Technology Jobs

Find similar Staff Software Engineer, Data Warehouse jobs: