Archer is best known for Midnight, our electric aircraft. This role isn't on that program. The AI Products Org is building software for the general aviation industry. As a Staff Data Engineer on the AI Platform team, you will design, build, and operate the data infrastructure that powers large-scale model training and inference. You will own the pipelines, storage systems, and data quality mechanisms that sit upstream of our ML platform, ensuring models train on clean, high-throughput, well-governed data.
What You'll Do:- Design and maintain high-throughput, fault-tolerant ingestion and transformation pipelines that feed training workloads at scale, with a focus on latency, throughput, and correctness.
- Build and operate the data lakehouse - defining table formats (Iceberg, Paimon, Parquet), partitioning strategies, and compaction policies optimized for ML consumption patterns.
- Instrument pipelines with data quality checks, lineage tracking, and anomaly detection so that model failures trace back to data problems quickly.
- Partner with ML engineers to define feature stores, dataset versioning, and experiment-to-production data contracts; integrate with tools like MLflow for dataset and artifact tracking.
- Work closely with AI researchers, platform engineers, and software engineers to understand data access patterns, optimize query performance, and unblock training runs.
What You Need:- 5+ years of professional data engineering experience excluding internships.
- BS/MS/PhD in Computer Science, Data Engineering, Software Engineering, or a related field.
- Hands-on experience building production pipelines with tools like Apache Spark, Flink, Airflow, dbt, or similar batch/streaming frameworks.
- Deep proficiency with columnar formats (i.e. Parquet), open table formats (Apache Iceberg or Apache Paimon), and object storage systems (S3 or equivalent).
- Experience with high-throughput messaging systems such as Apache Pulsar or Kafka for real-time data ingestion.
- Strong SQL skills; experience with StarRocks for large-scale analytical queries and real-time analytics over the lakehouse.
- Familiarity with AWS data services (S3, Glue, EMR) and containerized workloads (Docker/Kubernetes) in production environments. Airflow/Prefect/Dagster
Bonus Qualifications:- Experience with CDC (Change Data Capture) replication from transactional systems (i.e. PostgreSQL 12 Lakehouse via Debezium or Airbyte).
- Exposure to audio or time-series data pipelines, including preprocessing for ASR or speech model training.
- Prior experience in aerospace, aviation, or other safety-critical domains where data lineage and auditability are non-negotiable.
About The Team:AI Products is a roughly 100-person software org inside a company of aerospace engineers, and we expect to grow substantially over the next year. You get the resources and momentum of a real org, but the products themselves are early - which means Data Engineers here get the kind of ownership that usually only exists at startups.
If you want to hear more about what we're building before you apply, you can't. Apply, and we'll tell you plenty.
At Archer we aim to attract, retain, and motivate talent that possess the skills and leadership necessary to grow our business. We drive a pay-for-performance culture and reward performance that supports the Company's business strategy. For this position we are targeting a base pay between $175,00.00 - $215,000.00 . Actual compensation offered will be determined by factors such as job-related knowledge, skills, and experience.