Member of Technical Staff - Data Ingestion & Quality

Causal Labs

$120K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in building large-scale data pipelines and QA systems
  • Strong expertise in technologies like Apache Spark, Ray, or Beam
  • Excellent attention to detail for identifying subtle data inconsistencies
  • Ability to navigate and understand complex documentation for data ingestion
  • Proven experience working with external data vendors and managing partnerships

Responsibilities

  • Research and source new modalities of multimodal physical data and secure access through strategic partnerships
  • Build and maintain petabyte-scale data pipelines to ensure standardized data ingestion
  • Develop quality metrics to ensure data coverage, correctness, and consistency
  • Design automated QA checks to monitor data quality continuously and generate actionable insights
  • Write detailed technical requirements for external data vendors and provide constructive feedback
  • Collaborate with researchers to ensure dataset improvements enhance model performance

Benefits

  • Flexible work hours and remote work options
  • Opportunities for professional development and continuous learning
  • Collaborative and innovative work environment
  • Access to cutting-edge technology and resources
  • Health and wellness programs
  • Comprehensive health insurance packages
Full Job Description
Responsibilities

Your mission is to own every dataset end to end - from discovering the source and securing access, to writing the pipelines that ingest it, to guaranteeing it enters training clean, standardized, and correct.
  • Research and source new modalities of multimodal physical data (e.g. sparse sensors, point clouds, hyperspectral imagery, radar), and secure access through partnerships, vendors, and public archives
  • Build petabyte-scale data pipelines (e.g. Apache Spark) that ingest each source into our storage in standardized, training-ready form, across both batch and streaming - including the orchestration, storage, and monitoring they need where shared platform infrastructure doesn't yet exist
  • Develop quality metrics that measure coverage, correctness, and consistency across sources - and catch the subtle inconsistencies (sensor bias, drift, processing artifacts) that silently degrade models
  • Design and implement automated QA checks that continuously measure and monitor data quality over time, and own the verdicts they produce
  • Write technical requirements and provide actionable feedback to external data vendors and partners
  • Collaborate with researchers to validate that new and improved datasets translate into model performance


What we're looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
  • Demonstrated experience building large-scale data pipelines, QA systems, or evaluation workflows (e.g. Spark, Ray, Beam)
  • Detail-oriented in identifying subtle data inconsistencies and issues that could affect quality, with the ability to understand how quality impacts model performance
  • Comfortable going deep on unfamiliar source material - reading format specifications, sensor documentation, and vendor manuals to get ingestion exactly right
  • Experience working with external data vendors and partners, from technical evaluation to ongoing feedback
  • Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution

Similar Jobs

More Jobs at Causal Labs

More Information Technology Jobs

Find similar Member of Technical Staff - Data Ingestion & Quality jobs: