Staff Data Engineer

Harbor Compliance

$172K — $215K *
US-AnywhereRemote in United States
Enterprise Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 7+ years of hands-on data engineering experience, focusing on streaming and event-driven systems.
  • Proven experience designing and building near real-time data pipelines from scratch.
  • Hands-on production experience with vector databases and embeddings, ideally built from the ground up.
  • Advanced proficiency in SQL and Python.
  • Working knowledge of cloud warehouse/lakehouse platforms like Snowflake or BigQuery.
  • Experience contributing to production data environments, ideally as a founding data hire.
  • Familiarity with B2B SaaS and recurring revenue data models.

Responsibilities

  • Design, build, and own near real-time data pipelines as the backbone of the platform's data flow.
  • Evaluate, implement, and maintain vector database infrastructure for AI-augmented use cases.
  • Design and build ELT/ETL pipelines to ingest data from multiple sources, including HubSpot and HRIS.
  • Partner with leadership to architect the underlying data storage systems.
  • Build transformation layers to enable the Analytics Engineering team to create business-ready datasets.
  • Own pipeline reliability, including monitoring and automated alerting.
  • Build the foundation for self-service and AI-powered reporting for stakeholders.

Benefits

  • Health benefits
  • Flexible paid time off
  • Parental leave
  • Fertility and adoption assistance
  • 401(k)
  • Educational reimbursement
Full Job Description
Staff Data Engineer - unifying fragmented data from HubSpot, financial systems, and HRIS into a single, AI-augmented source of truth for executive decision-making. As Staff Data Engineer, you'll be a foundational technical hire, working closely with the Sr. Director of Data Platform & Analytics to design and build the real-time, AI-ready data infrastructure that powers this platform. This is a hands-on role for someone who wants to build from the ground up, not maintain what already exists.

Key Responsibilities

  • Design, build, and own near real-time data pipelines (CDC, streaming ingestion, event-driven architectures) as the backbone of the platform's data flow
  • Evaluate, implement, and maintain vector database infrastructure and embedding pipelines to support AI-augmented use cases (semantic search, retrieval-augmented generation, AI agents acting on company data).
  • Design and build ELT/ETL pipelines ingesting data from our platform, financial platforms, CRM (HubSpot) and HRIS, feeding both real-time and batch use cases.
  • Partner with the Sr. Director to architect the underlying warehouse/lakehouse as a supporting system of record - the storage layer beneath the streaming and AI infrastructure.
  • Build lightweight transformation layers (e.g., dbt) as needed to enable our Analytics Engineering team translate raw data into business-ready datasets aligned to core metrics like ARR, CAC, and churn.
  • Own pipeline reliability and observability - monitoring, automated failure alerting, and lineage tracking across both streaming and batch pipelines.
  • Build the technical foundation for self-service and AI-powered reporting, partnering with BI, Product & Engineering stakeholders on recurring executive and departmental reports.
  • Implement data governance practices, including documentation standards and access controls.
  • Partner cross-functionally with Finance, Marketing, Customer Success, and Operations to translate data needs into reliable, low-latency data products.
  • Leverage AI-augmented development workflows (e.g., Claude Code) to accelerate pipeline development and documentation.

Requirements

  • 7+ years of hands-on data engineering experience, with meaningful depth in streaming/event-driven systems, not just batch pipelines.
  • Proven experience designing and building near real-time pipelines from scratch (e.g., Kafka, Kinesis, Flink, Debezium/CDC) in a production environment.
  • Hands-on production experience with vector databases and embeddings (e.g., Zilliz, Pinecone, Weaviate, pgvector, Milvus) - ideally having built this infrastructure from the ground up rather than just consuming a managed AI feature.
  • Advanced proficiency in SQL and Python.
  • Working knowledge of a cloud warehouse/lakehouse platform (Snowflake, BigQuery, or Databricks) and dbt - you'll use these, but they're the storage/transform layer supporting the streaming and AI work, not the main focus.
  • Proven experience building or materially contributing to an end-to-end production data environment, ideally as an early or founding data hire.
  • Familiarity with B2B SaaS and recurring revenue data models (customer lifecycle, pipeline/conversion data).
  • Working knowledge of BI/reporting tools (e.g., Looker, Tableau, Power BI) as a downstream consumer of your data models.
  • Ability to work independently and drive multi-stakeholder projects in a lean, scrappy, fast-moving environment - comfortable with ambiguity and building without a lot of existing infrastructure or process.

Skills and Knowledge

  • Strong command of event streaming and CDC tooling.
  • Hands-on experience with vector databases (Pinecone, Weaviate, pgvector, Milvus, or similar), including embedding strategies and chunking approaches for retrieval use cases.
  • Working knowledge of ELT/ETL tools (e.g., Fivetran, Airbyte) for batch use cases.
  • Working knowledge of data observability/reliability practices - automated alerting, lineage tracking.
  • Fluency with AI-augmented development tools (e.g., Claude, Copilot) to accelerate engineering and documentation.


Compensation:

Harbor Compliance's base salary range for this role is listed below. Compensation at the time of offer is based on factors such as skill set, experience, qualifications, and work location. Salary is one part of Harbor Compliance's total compensation package. Other benefits may include health benefits, flexible paid time off, parental leave, fertility and adoption assistance, 401(k), and educational reimbursement. Note that the salary range and benefits apply only to U.S.-based candidates.

The pay range for this role is:

172,000 - 215,000 USD per year (US National)

Similar Jobs

More Jobs at Harbor Compliance

More Enterprise Technology Jobs

Find similar Staff Data Engineer jobs: