Job Summary
The Lead Data Engineer will provide technical leadership and mentorship to a data engineering team while remaining primarily hands-on in development. The role will focus on building scalable and maintainable data solutions, ensuring data quality, security, privacy, and integrity throughout the data lifecycle, and providing technical guidance across application architecture and development teams. The position will support new initiatives across AI, Sales, Marketing, and Customer Analytics, with an emphasis on GCP, SQL, Python, distributed data processing, and data engineering best practices.
Key Responsibilities
• Mentor data engineers and provide technical guidance, best practices, and support to help the team meet project deadlines and priorities.
• Provide proactive technical oversight and advice to application architecture and development teams.
• Promote re-use, scalability, stability, and operational efficiency of data and analytical solutions.
• Ensure the quality, completeness, security, privacy, and integrity of data throughout the data lifecycle.
• Develop and maintain a deep understanding of data sources, granularity, availability, and limitations.
• Create maintainable and scalable code to load and manipulate data in data warehouses.
• Develop ETL workflows and data-driven solutions using Python.
• Design and implement distributed data processing solutions using big data technologies such as Hadoop or Spark.
• Apply design patterns and develop generalized, reusable code for common use cases.
• Contribute to reusable libraries and maintain high-quality engineering standards.
• Document critical workflows and operational support aspects of the team's responsibilities.
• Facilitate communication across project teams, business stakeholders, and technical teams.
• Develop and consume web services to support data and analytical solutions.
• Support new projects involving AI, Sales, Marketing, and Customer Analytics.
• Maintain a primarily hands-on development focus while providing mentorship and technical leadership.
Required Qualifications
• 8-10 years of experience in data engineering or a related technical discipline.
• Expert-level experience with Google Cloud Platform (GCP).
• Expert-level SQL experience.
• Extensive hands-on Python programming experience, particularly for ETL workflows and data-driven solutions.
• In-depth knowledge of SQL or NoSQL and experience working with various data stores, including relational databases, analytical databases, and scalable document stores.
• Expertise with big data batch computing technologies such as Hadoop or Spark, including distributed data processing.
• Applied knowledge of data modeling principles, including dimensional modeling and star schemas.
• Strong understanding of database internals, including indexes, binary logging, and transactions.
• Experience with infrastructure-as-code and related technologies such as Docker, CloudFormation, or Terraform.
• Experience with software engineering tools and workflows, including Git, Jenkins, and CI/CD.
• Practical experience authoring and consuming web services.
• Ability to develop robust, high-quality, reusable code.
• Demonstrated experience providing customer-driven solutions, support, or service.
• Strong communication skills with the ability to work effectively with business stakeholders.
• Solid data understanding and business acumen, particularly in data-rich industries such as insurance or financial services.
Preferred Qualifications
• Knowledge of open-source machine learning toolkits such as scikit-learn, Spark ML, or H2O.
• Working knowledge of distributed systems in cloud environments.
• Working knowledge of machine learning engineering.
• Experience supporting AI, Sales, Marketing, or Customer Analytics initiatives.
• Experience working in business-facing technical roles.