4-6+ years of software/data engineering experience, particularly with large-scale data platforms.
Deep knowledge of lakehouse architectures and technologies like Apache Iceberg and Delta Lake.
Hands-on experience with vector databases and embedding retrieval systems.
Strong familiarity with AWS or other major cloud providers for data services and cost optimization.
Proficiency in Python and SQL, with experience in orchestration tools like Airflow or Prefect.
Experience designing data architectures for ML/AI workloads, including training data pipelines.
Solid understanding of data modeling and the trade-offs between batch and streaming.
Responsibilities
Architect and build a data lakehouse, including ingestion pipelines and governance.
Design and operate vector database infrastructure for large-scale embedding storage and retrieval.
Build large-scale data management systems for lifecycle, lineage, and quality monitoring.
Create data platform capabilities that support ML use cases, including feature pipelines and low-latency retrieval.
Establish schema and data-contract standards, enabling self-service for researchers and engineers.
Ensure reliability and performance of data pipelines in production, focusing on observability.
Mentor engineers and lead technical design across the data domain.
Benefits
Above-market salary, equity, and benefits package.
Early Series A equity.
Excellent health, dental, and vision coverage.
401(k) match up to 4% of your salary.
Flexible PTO.
Daily office lunches in NYC.
Full Job Description
The Role
DS collects radio data at a scale few organizations ever see - continuous, high-rate streams from sensors around the globe. We're hiring a Senior Data Engineer to design the platform that turns that firehose into an asset: a data lakehouse that serves researchers training models, production systems running inference, and agents retrieving context in real time.
You'll own the architecture from ingestion through storage, cataloging, vector search, and access, and you'll design it explicitly for AI/ML workloads, not just analytics. What you'll do
Architect and build our data lakehouse: ingestion pipelines, open table formats, partitioning and compaction strategies, cataloging, and governance.
Design and operate vector database infrastructure for embedding storage, similarity search, and retrieval at scale - choosing, tuning, and evolving the right systems for our workloads.
Build large-scale data management systems: lifecycle and retention, lineage, versioning of datasets for reproducible training, quality monitoring, and cost management across petabyte-class storage.
Design data platform capabilities that directly support ML use cases - feature and embedding pipelines, training-set assembly, evaluation datasets, and low-latency retrieval for agents.
Establish schema and data-contract standards across teams, and build tooling that lets researchers and engineers self-serve.
Own reliability and performance of data pipelines in production, including observability and failure handling.
Mentor engineers and lead technical design across the data domain.
What we're looking for
4-6+ years of software / data engineering experience, including several years designing and operating large-scale data platforms in production.
Deep experience with lakehouse architectures and technologies - e.g., Apache Iceberg, Delta Lake, or Hudi; Parquet; Spark, Flink, Trino, DuckDB, or similar engines.
Hands-on experience with vector databases and embedding retrieval (e.g., pgvector, Milvus, Qdrant, Weaviate, Pinecone, LanceDB, or FAISS-based systems), including indexing trade-offs and scaling.
Strong experience with AWS or another major cloud provider - object storage, managed data services, compute, and cost optimization at scale.
Fluency in Python and SQL; experience with orchestration tools (Airflow, Dagster, Prefect, Step Functions, etc.).
Demonstrated experience designing data architectures specifically for ML / AI workloads: training data pipelines, feature stores, dataset versioning, or retrieval systems.
Strong grasp of data modeling, consistency, and the trade-offs between batch and streaming.
Nice to have
Experience with streaming ingestion at high volume (Kafka, Kinesis, Pulsar).
Experience with time-series, geospatial, or signal / sensor data.
Familiarity with data governance and security requirements in regulated environments.
Experience with Rust, Go, or C++ for performance-critical data paths.
Who Thrives at Distributed Spectrum
Fast learners over specific backgrounds - We care more about how quickly you can pick up new skills than where you've worked before.
Intellectual honesty - The right answer matters more than being right. You challenge assumptions, test ideas, and pivot when needed.
Adaptability - We're organized, but sometimes things change quickly. You find a way to make it work and balance short-term deliverables with long-term goals.
Ownership of outcomes - You optimize your own time, focus on what matters to deliver quickly, and cut out inefficiencies.
Not building in a vacuum - You stay connected to the rest of our teams and our customers to make sure all the pieces fit together.
What We Offer
Above-market salary, equity, and benefits package.