Member of Technical Staff - Data Infrastructure

Causal Labs

$130K — $180K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in data engineering or a related field
  • Proven track record building and maintaining large-scale data pipelines
  • Expertise in distributed computation tools like Spark, Ray, or Beam
  • Familiarity with cutting-edge data storage solutions and file formats such as Parquet and Delta Lake
  • Strong knowledge of cloud infrastructure and data lake architectures
  • Ability to drive projects independently from inception to delivery

Responsibilities

  • Design and manage petabyte-scale data storage systems
  • Oversee a shared compute and orchestration platform for data processing
  • Enhance data strategy from storage through to loading for efficient training
  • Develop systems for data cataloging, deduplication, and lineage tracking
  • Implement quality and monitoring tools for data integrity
  • Scale infrastructure to boost engineering efficiency and reliability
  • Engage in all phases of the data lifecycle, including ingestion of critical data

Benefits

  • Opportunities for professional growth and development
  • Access to the latest technology and resources
  • Flexible work environment that fosters creativity
  • Collaborative teamwork culture that encourages innovative solutions
  • Health and wellness programs to support work-life balance
Full Job Description
We look for data engineers who are excited to tackle unsolved problems. Physical observations arrive continuously, in many formats, at a scale that dwarfs what is used to train today's LLMs. Your mission is to build the data platform underneath it all - the storage, compute, and loading systems that make every dataset cheap to ingest, fast to query, and immediately available to training.

Responsibilities
  • Design and operate petabyte-scale storage: lakehouse architecture, file formats, and data layout optimized for both batch and real-time queries
  • Own the shared compute and orchestration platform (e.g. Spark, Ray, workflow scheduling) that ingestion and research pipelines run on
  • Optimize data strategy end to end from storage to loading, owning high-throughput data loading into training up to the tensor boundary
  • Build systems for cataloging, deduplication, lineage, search, and reproducibility at every stage of the data lifecycle
  • Implement the platform-level quality and monitoring tooling that data and research teams build their checks on
  • Scale infrastructure to improve engineering velocity and ensure reliability, with monitoring and alerting to match
  • Work across the full data lifecycle when the mission needs it - including building and operating ingestion pipelines for critical data sources directly


What we're looking for

We value a relentless approach to problem-solving, rapid execution, and the ability to quickly learn in unfamiliar domains.
  • Demonstrated experience building large-scale data pipelines and distributed compute systems (e.g. Spark, Ray, Beam)
  • Knowledge of state-of-the-art methods and tools for data ingestion, storage, and loading - including file formats and storage systems (e.g. Parquet, Zarr, Delta Lake) and how they impact performance and scalability
  • Deep familiarity with cloud infrastructure, data lake architectures, and batch and streaming pipelines
  • Understanding of how data loading throughput affects large-scale training, and experience optimizing it
  • Owns deliverables end-to-end, from collecting and translating requirements to autonomously driving execution

Similar Jobs

More Jobs at Causal Labs

More Information Technology Jobs

Find similar Member of Technical Staff - Data Infrastructure jobs: