Your role and responsibilities:
As a Data Platform Engineer, you will take Xanadu's existing data pipelines and transform them into a cohesive data platform. Working alongside hardware researchers, data scientists, and finance analysts, you'll define how data flows across the organization and build the infrastructure that enables data-driven decisions at every level. This is a hands-on technical role first - you will be the first dedicated data engineer, with the opportunity to build a data team around you.
This is a high-complexity, moderate-volume problem. Terabytes, not petabytes. Different heterogeneous measurement types, not billions of uniform events.
Your responsibilities will be:
- Design, build and maintain a robust cloud data infrastructure (ingestion, transformation, serving) that will serve multiple engineering team across the organization
- Own the data architecture: layering, schema and contract design, materialization strategy, query performance.
- Understand and consolidate existing data workflows and pipelines spanning R&D, manufacturing, and business analytics
- Define data models in collaboration with researchers and analysts to ensure scientific and business data is stored correctly and queryable
- Establish data governance foundations: lineage, cataloging, access control, and quality monitoring
- Drive best practices across code quality, testing, data reliability, and observability
- Balance new development, platform improvements, and technical debt reduction
- Mentor and provide technical guidance to engineers across the organization as the practice grows
Basic qualifications and experience:
- 7+ years in data engineering, with 2+ years in a lead or architect capacity
- Deep experience building and scaling data platforms in the cloud from the ground up
- Strong software engineering: Python packaging, testing, CI
- Production experience designing and operating data lake or lakehouse architectures (Delta Lake, Iceberg, or Hudi)
- Hands-on experience with modern data stack tooling (dbt or similar) and orchestration (Airflow, Dagster, Prefect)
- Strong SQL skills
- Knowledge of infrastructure-as-code and CI/CD for data pipelines
- Proven ability to drive technical standards and engineering improvements across teams
- Experience working with cross-functional teams - especially R&D or science teams producing unstructured or semi-structured data
Preferred qualifications and experience
- Experience in a deep-tech, hardware, or semiconductor environment where data originates from physical measurement systems
- Familiarity with time-series or scientific data formats (HDF5, Parquet for measurement traces, etc.)
- Prior experience as the "first data engineer" - building a practice from scratch, not inheriting one
- Familiarity with LIMS systems or laboratory data workflows
This is for a new position. Your base salary will be determined based on your location, experience, and internal benchmarks. You will also be eligible for equity and benefits.