Full Job Description
Lead Data Software Engineer with Databricks
We are seeking a Lead Data Software Engineer with Databricks expertise to drive the technical design, robustness, and evolution of end-to-end traceability solutions, enabling transparency, regulatory compliance, and data-driven decision-making across the supply chain. This role provides technical leadership and architectural direction for traceability platforms and data products, working with heterogeneous datasets ranging from farm and geolocation data to transactional, logistics, sustainability, and regulatory information. The successful candidate will collaborate closely with the Traceability solution team, sustainability experts, business analysts, data engineers, and data scientists, acting as a technical reference to ensure traceability data pipelines, data models, and platforms remain aligned with best practices and long-term architectural vision.
Responsibilities
Provide technical leadership and ownership for traceability data platforms, pipelines, and data products
Define, document, and enforce architectural principles, technical standards, and best practices for traceability solutions
Oversee the design of robust pipelines, quality controls, and proactive monitoring frameworks to ensure data availability, quality, integrity, and timeliness
Guide and support data engineers and technical contributors on implementation choices, troubleshooting, performance optimization, and best practices
Contribute strategically to Run execution and continuous improvement, including data quality, lineage, auditability, and operational monitoring
Translate traceability and regulatory requirements into technical solutions, data models, and platform capabilities
Collaborate closely with solution architects, data platform teams, and external partners to ensure alignment and scalability
Support adoption of traceability solutions by business and sustainability teams through technical guidance and enablement
Requirements
5+ years of experience in technical roles within data, analytics, or platform environments, with exposure to leadership or solution ownership responsibilities
Expertise in cloud-based data platforms, preferably Azure
Background in data architecture, data modeling, and data platform design, including complex, large-scale datasets
Hands-on experience with Databricks, Spark/PySpark, and SQL for data transformation pipelines and analytics platforms
Proficiency in SQL and data storage concepts, including relational databases, NoSQL, and data lakes, with performance optimization for large and complex datasets
Knowledge of Python and data manipulation ecosystems such as Pandas and NumPy, with ability to review and guide code even if not hands-on full time
Familiarity with data visualization tools including Power BI, Tableau, and Looker
Experience designing or governing data quality frameworks, lineage, monitoring, and auditability
Familiarity with orchestration tools and data integration patterns such as ADF, Airflow, and AWS Glue, or equivalent
Experience with software engineering best practices including version control with git, branching strategies, code reviews, and CI/CD
Understanding of data governance, security, and access control for sensitive business and regulatory data
Proficiency in English at a B2+ level