Full Job Description
Lead Data Software Engineer
We're looking for a hands-on, experienced Lead Data Software Engineer who has a deep understanding of Apache Spark and large-scale distributed computing. This role covers the end-to-end lifecycle of our data infrastructure, from designing and implementing ingestion solutions to ensuring the long-term reliability and scalability of our data pipelines. The ideal candidate is a collaborative professional who can work with cross-functional teams, translate complex business requirements into high-quality technical solutions, and maintain a rigorous focus on engineering standards and performance.
Responsibilities
Design and implementation of robust data ingestion solutions and pipelines using Cloud Native and Big Data technologies, including architecting key system components and selecting the appropriate tech stack to ensure long-term scalability, performance, and alignment with the broader technical roadmap
Team leadership through daily guidance and mentorship so the team delivers high-quality, scalable solutions on schedule, along with serving as the primary technical point of contact for customers, translating business visions into actionable technical roadmaps and managing expectations around delivery, feasibility, and progress
Full technical stack management, including configuration management, monitoring, debugging, and performance tuning of data solutions
Development and maintenance of scalable data pipelines for efficient data processing and analysis
Partnership with cross-functional teams and data architects to identify business requirements and translate them into technical specifications
Code reviews and development testing leadership and participation to ensure adherence to standards, quality gates, and SDLC best practices
Creation and maintenance of detailed technical documentation for all data engineering projects to support transparency and knowledge sharing
Timely and efficient troubleshooting and resolution of data-related issues, ensuring all solutions remain high-quality, reliable, and scalable
Requirements
5+ years of experience in Data Software Engineering with a focus on large-scale distributed systems
A demonstrated track record of designing modular system components and selecting technology stacks (such as storage formats and processing engines) that ensure long-term performance, scalability, and alignment with the broader technical roadmap
Experience as a Team Lead providing technical guidance to ensure high-quality delivery, while acting as a technical voice for customers to bridge the gap between business needs and engineering solutions
Proficiency in Python and SQL
Extensive hands-on experience with Apache Spark, Kafka, and Airflow
Proficiency with major cloud providers (AWS, Azure, or GCP)
Knowledge of Databricks or Snowflake (nice to have)
Practical experience in Apache Spark job performance tuning
Strong understanding of the SDLC, quality gates, and Agile methodologies
Excellent communication skills with English proficiency at a B2+ level
Motivation, independence, and the ability to handle several projects simultaneously with a focus on efficiency