Software Engineer, Data Ingestion Systems

Wayve

• $120K — $145K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of production experience with Apache Spark and Python
  • Experience building and debugging large-scale ingestion or ETL pipelines
  • Ability to diagnose production failures involving corrupt data and complex dependencies
  • Understanding of optimizing high-throughput systems for reliability and cost
  • Quick learner capable of working independently in a fast-paced environment
  • Familiarity with Airflow, Flyte, Databricks, Delta Lake, Scala, or Java is a plus

Responsibilities

  • Debug and resolve issues in ingestion pipelines and processing queues
  • Investigate and rectify corrupt or malformed data from various sources
  • Optimize Spark jobs and batch-processing pipelines for efficiency
  • Collaborate with engineers to prioritize critical datasets and reduce operational challenges
  • Enhance monitoring and recovery processes across multi-step ingestion workflows

Benefits

  • Hybrid working model with flexible remote options
  • Learning and development budgets for training and conferences
  • Comprehensive health and dental insurance
  • Enhanced maternity and paternity leave
  • Relocation support and visa sponsorship where applicable
  • Opportunities for hands-on work in vehicle workshops and labs
  • Meaningful equity in the company
Full Job Description
Before the detail, here's the challenge you'd help us solve.

We build the embodied intelligence that moves real vehicles safely, and the ecosystem a billion machines will run on in the future. Very few people in AI can say this. Every role here, whatever the team, plugs into that.

Here's what this particular role covers.

About our Ingestion Systems Team

Our Ingestion Systems team sits within Wayve's AI Platform organisation and builds the data foundations behind our self-driving technology. We process vast amounts of real-world driving data so it can be used for annotation, model training and evaluation. Operating at petabyte scale, the team focuses on making these critical data flows reliable, efficient and easy to operate.

Your day-to-day
  • Debug and unblock failing ingestion pipelines, stalled jobs and processing queues
  • Investigate corrupt, malformed or unexpected data from a variety of sources.
  • Optimise Spark jobs and batch-processing pipelines for throughput, reliability and compute efficiency.
  • Work alongside permanent engineers to prioritise critical datasets and reduce operational toil.
  • Improve monitoring, retries and failure recovery across multi-step ingestion workflows.

What you'll be working on
  • High-volume pipelines that process real-world driving data for annotation, data science, model training and evaluation.
  • Resilient systems that isolate problematic data rather than allowing a single bad segment to block a wider workflow.
  • Orchestration across complex dependencies, queues, retries and priority data flows.
  • Pragmatic improvements to the performance, cost and operability of our Spark-based data platform.
  • Ingestion of varied data formats from partners, suppliers and other third-party sources.

You should apply if
  • You have strong production experience with Apache Spark and Python.
  • You have built, operated or debugged large-scale ingestion, ETL or distributed data-processing pipelines.
  • You are comfortable diagnosing production failures involving corrupt data, stalled queues, retries and complex dependencies.
  • You understand how to optimise high-throughput systems for reliability, performance and cost.
  • You can learn an existing system quickly, work independently and deliver practical improvements in a fast-moving environment.
  • Experience with Airflow, Flyte, Databricks, Delta Lake, Scala or Java would be an advantage.


Not ticking every box? That's totally okay! If you're passionate about autonomy and keen to learn, we encourage you to apply even if you don't meet every requirement.

How we work - Locations & Flexible Working:

Our main hubs are in London, Sunnyvale, Yokohama, Herzliya, Vancouver and Leonberg. We operate a hybrid working model that combines in-person collaboration in our dedicated office spaces with focused time working remotely. This gives our teams the connection and energy of working together, alongside the flexibility to do their best work in a way that fits their lives.

The Interview Process

Our process is clear and respectful of your time
• Initial call / recruiter screen (30 mins)
• Competency interview Python Programming (60 mins)
• Deep-dive technical interview: SQL Domain (60 mins)
• Final interview: Mission and values alignment (45 mins)

We'll always explain the format and work around your availability.

What's in it for you (Location dependant):

Salaries benchmarked against the market annually
Meaningful equity, sharing in the ownership and long term success of Wayve
Relocation support and visa sponsorship where applicable
• Hybrid working, core hours and the chance to work hands on in vehicle workshops and labs
Learning and development budgets with support for training, conferences and growth
Comprehensive benefits including health insurance, dental, enhanced maternity and paternity leave, retirement or pension where applicable, access to therapists, wellbeing partnerships, team socials and more

A quick, honest note before you apply.

Wayve is not a mature, fully-structured place with the playbook already written. Much of how we work is still being written, and if you join, you'll help write it. That suits people who want real ownership more than people who need a settled structure from day one.

If that sounds like the kind of problem you want to spend your time on, we'd really like to hear from you.

Similar Jobs

More Jobs at Wayve

More Information Technology Jobs

Find similar Software Engineer, Data Ingestion Systems jobs: