Lead Data Engineer

Compunnel

$120K — $145K *
Finance & Insurance
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's in Computer Science, Data Engineering, or related field.
  • 5+ years in Data Engineering or Data Platform development.
  • Experience with financial or payroll data structures.
  • Proficiency in Databricks and Delta Lake with a strong understanding of Unity Catalog governance.
  • Strong Python, SQL, and PySpark skills with experience in ETL/ELT development.

Responsibilities

  • Design and maintain scalable data pipelines for various financial datasets.
  • Build Databricks platforms to support financial analytics.
  • Implement ETL/ELT frameworks using PySpark for diverse data sources.
  • Develop optimized data models for analytical and reporting use cases.
  • Ensure data quality and governance across data assets.

Benefits

  • Opportunities to work with cutting-edge technologies like AI-assisted coding tools.
  • Collaboration with cross-functional teams like data scientists and economists.
  • Involvement in design discussions for advanced data architecture.
Full Job Description
Job Summary:

The Lead Data Engineer will work across data engineering and analytical responsibilities, with a primary focus on building scalable cloud-native data platforms and supporting financial analytics and research. The role requires strong expertise in Databricks, PySpark, Python, SQL, Delta Lake, data modeling, and large-scale data processing, along with finance or payroll domain experience. The engineer will collaborate closely with data scientists, economists, and business stakeholders to develop production ETL pipelines, explore and profile datasets, optimize data platforms, and translate complex analytical requirements into scalable solutions.

Key Responsibilities
• Design, develop, and maintain scalable data pipelines for the ingestion, transformation, and distribution of payroll, macroeconomic, and financial datasets.
• Build and support Databricks-based platforms that enable financial research and analytical workloads.
• Implement ETL/ELT frameworks using PySpark and Delta Lake for structured and unstructured data from internal and external sources.
• Develop data models and data marts optimized for analytical, reporting, and machine learning use cases.
• Ensure data quality, consistency, lineage, governance, and observability across data assets.
• Optimize performance for large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
• Translate business requirements into scalable and maintainable data solutions.
• Perform exploratory data analysis, including dataset profiling and identification of distributions, outliers, missing patterns, and data drift.
• Translate data scientist logic into efficient and scalable PySpark implementations, including cross-sectional metrics and time-windowed aggregations.
• Build validation dashboards and exploratory notebooks to verify pipeline outputs and data quality.
• Support feature engineering by implementing complex aggregation and transformation logic at scale.
• Independently validate analytical outputs, perform sanity checks, and identify results that warrant further investigation.
• Conduct ad hoc analytical work using pandas and NumPy alongside PySpark to support research initiatives.
• Contribute to lakehouse architecture design discussions and evaluate tradeoffs related to catalog design, medallion architecture, and data mesh concepts.
• Implement CI/CD pipelines using Databricks Asset Bundles, Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
• Manage Unity Catalog governance, access patterns, and schema design.
• Ensure security, compliance, and data governance standards are maintained.
• Leverage AI-assisted coding tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent platforms to accelerate development.
• Review AI-generated code for correctness, performance, scalability, and maintainability.
• Integrate AI-assisted development workflows into engineering and analytical activities.

Required Qualifications
• Bachelor's or Master's degree in Computer Science, Data Engineering, Information Systems, Statistics, Economics, Finance, or a related field.
• 5+ years of experience in Data Engineering or Data Platform development.
• Finance or payroll domain experience, including familiarity with payroll data structures, pay-period logic, compensation and deduction relationships, or financial-services data environments.
• Experience handling large-scale datasets, including billions of records, multi-terabyte environments, and time-series data.
• Strong proficiency with Databricks, including Unity Catalog, Delta Lake, Databricks Workflows, and Databricks Asset Bundles or equivalent deployment frameworks.
• Strong understanding of Unity Catalog governance, access patterns, and catalog/schema design.
• Strong understanding of Delta Lake internals, including optimization, clustering, change data feed, and versioning.
• Experience with Databricks Workflows, including orchestration, dependencies, and monitoring.
• Experience participating in architecture-level design decisions and evaluating technical tradeoffs.
• Strong proficiency in Python, SQL, PySpark, data modeling, and ETL/ELT development.
• Analytical fluency with exploratory data analysis, basic statistical concepts, distributions, correlations, time-series patterns, and feature engineering.
• Proficiency with pandas and NumPy for ad hoc analytical work alongside production PySpark.
• Experience implementing CI/CD using tools such as Bitbucket Pipelines, Jenkins, and automated deployment frameworks.
• Experience with AI-assisted development tools such as GitHub Copilot, Amazon Q, Kiro, or equivalent.
• Experience implementing data quality and validation frameworks.
• Strong understanding of data governance, data quality, and metadata management.
• Strong analytical and problem-solving skills with statistical literacy.
• Excellent communication and documentation skills.
• Ability to work effectively in a fast-paced, data-driven environment and collaborate with technical and business stakeholders.

Preferred Qualifications
• Experience with macroeconomic, capital markets, or financial-services data.
• Exposure to lakehouse patterns, data mesh concepts, and medallion architecture.
• Experience with streaming and event-driven pipelines using Kafka or Structured Streaming.
• Census or geographic data processing experience, including TIGER and FIPS codes.
• Infrastructure-as-code experience with Terraform, CDK, or similar technologies.
• Databricks Associate or Professional-level certification.
• Experience migrating legacy data platforms to modern technology stacks, including Glue, EMR, or HDInsight to Databricks.
• Experience supporting machine learning and AI-driven analytics solutions.
• Data visualization experience using Power BI, Tableau, Databricks Dashboards, or similar platforms.
• Experience with Python visualization libraries such as Matplotlib and Plotly.
• Experience with SQL Server, PostgreSQL, Delta Tables, or NoSQL databases.
• Experience with Scala.
• Experience with large-scale time-series data engineering.
• Experience with payroll data structures and processing.
• Experience with macroeconomic analysis and forecasting.
• Experience with financial markets and alternative data.

Similar Jobs

More Jobs at Compunnel

More Finance & Insurance Jobs

Find similar Lead Data Engineer jobs: