The roleWe are hiring a senior Software Engineer to build the platform that powers Wayve's data flywheel and foundation-model stack. This is the
engineering counterpart to our Applied Scientist and ML Engineer roles: you build the systems they, and the wider company, depend on. It is high-leverage, high-visibility work with a clear path to deep system ownership.
- Build the systems that allow teams to turn world-scale driving data into high-signal training data, and evaluate and train foundation models on it.
- Replace ad-hoc scripts and manual handoffs with self-serve, observable products used across Science, Autonomy, and Evaluation.
- Every model Wayve ships runs on this platform: your work compounds across the entire fleet and roadmap.
- Work shoulder to shoulder with a world-class science and engineering team, with real deployment at global OEM scale (Nissan, Stellantis, Uber).
- TC3 / TC4 ownership of platform and infrastructure, with room to set technical direction as the platform matures.
What you will do- Build and scale the data curation and enrichment pipelines that turn world-scale fleet data into high-signal training data: mining and active-learning loops, running model-based enrichments over billions of rows, and ensuring data quality at scale.
- Build the evaluation infrastructure behind foundation-model progress: harnesses for offline and closed-loop evaluation, metric and benchmark pipelines, and world-model-based evaluation.
- Build and optimize training and serving infrastructure for large pretrained models: distributed training, batched inference, and large-scale model backfills.
- Build the data-platform backbone: distributed data processing (Ray Data, Daft, Spark / Databricks), embedding and vector search (turbopuffer, Milvus), lakehouse formats (Lance, Iceberg), dataset versioning, and the enrichment and annotation catalog.
- Make it self-serve and reliable: turn one-off processes into products that other teams operate themselves, and own testing, observability, and on-call for what you ship.
- Partner closely with Applied Scientists and ML Engineers to take research from prototype to production at scale.
What we are looking for- Strong production software engineering, especially production Python (services, APIs, large-scale data processing), and comfort owning and extending large codebases.
- Large-scale data and distributed-systems experience: batch and streaming pipelines, workflow orchestration (Flyte, Airflow, Dagster, or similar), and distributed processing (Spark / PySpark, Ray, Databricks, or equivalent).
- Systems design for scale: reliable, observable, high-throughput data or ML systems, with strong SQL and query and performance optimization.
- A track record of shipping and operating production systems that other teams depend on: testing, code review, observability, and on-call.
- Strong CS fundamentals and several years of production experience (roughly 6 or more for TC3, more for TC4), or equivalent; a degree in CS or comparable practical experience.
- Seniority to match the level: takes ambiguous, cross-team problems and drives them to completion, and at TC4 sets technical direction and multiplies the team.
Bonus points- ML platform / MLOps: model registration, distributed training, and inference or serving optimization.
- Enough exposure to foundation models, world models, or ML evaluation to partner deeply with scientists.
- Embedding and vector search, annotation tooling, or feature and data catalogs.
- Kubernetes and modern data / lakehouse stacks (Databricks, Lance, Iceberg).
- Autonomous driving, robotics, or other large-scale sensor-data workflows.