Lead Data Engineer - PySpark/Palantir Foundry

Logic20/20 Inc.

• $156K — $175K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • 10-15+ years in data engineering, data science, or machine learning engineering with Python expertise.
  • Proven experience leading technical teams on enterprise-scale data projects.
  • Strong skills in PySpark, SQL, and cloud services.
  • Ability to refactor and stabilize existing codebases and pipeline environments.
  • Experience with cloud-optimized datasets and large-scale spatial operations.
  • Familiarity with machine learning model outputs and their integration into data pipelines.
  • Background in highly regulated industries like utilities, finance, or healthcare.

Responsibilities

  • Design and maintain production-grade data pipelines for model output aggregation.
  • Establish repository governance practices including branching and version control.
  • Own release engineering practices for reproducible and auditable production releases.
  • Refactor pipeline code to enhance modularity and maintainability.
  • Develop configuration-driven pipeline patterns for consistency across environments.
  • Support testing and validation for critical data pipelines.
  • Collaborate with cross-functional teams to align on data schemas and delivery expectations.

Benefits

  • Recognition and rewards for exceptional talent.
  • Competitive compensation package with performance-based bonuses.
  • Opportunities for professional growth and development.
Full Job Description
Job Description

As a Lead Data Engineer joining our AI & Analytics practice, you'll be responsible for delivering client value and ensuring high client satisfaction. You'll be expected to be adept at recognizing, subscribing, and applying best practices, methodologies, tools, and techniques to meet client requirements, timelines, and budgets.

The right candidate will be highly hands-on, comfortable working across cross-functional teams, and motivated by clean architecture, disciplined engineering practices, and well-governed delivery processes.

Role Responsibilities
  • Design, enhance, and maintain production-grade data pipelines that support model output aggregation and downstream risk analysis.
  • Establish and mature repository governance practices, including branching strategy, pull request standards, merge policies, release tagging, and version control workflows.
  • Own release engineering practices that support reproducible, traceable, and auditable production releases.
  • Refactor and improve existing pipeline code to increase modularity, maintainability, scalability, and documentation quality.
  • Develop configuration-driven pipeline patterns that create consistency and reduce ambiguity across environments and releases.
  • Support testing, validation, benchmarking, and change management practices for critical data pipelines.
  • Partner with data scientists, machine learning engineers, data engineers, product stakeholders, and other technical teams to align on schemas, interfaces, inputs, and delivery expectations.
  • Translate complex technical concepts into clear updates for both technical and non-technical stakeholders.
  • Help bring structure to existing codebases, repositories, and engineering workflows that may need stronger governance or standardization.
  • Contribute to engineering best practices across a highly regulated, audit-sensitive delivery environment.


Qualifications
  • 10-15+ years of data engineering, data science, machine learning engineering, and/or relevant experience using Python.
  • Experience leading technical teams and overseeing enterprise-scale data initiatives.
  • Strong expertise in PySpark, SQL, and cloud services.
  • Demonstrated ability to improve, refactor, or stabilize existing codebases and pipeline environments.
  • Experience with cloud-optimized datasets, efficient partitioning strategies, and large-scale spatial operations.
  • Strong understanding of how machine learning model outputs flow into downstream data pipelines, platforms, or production systems.
  • Experience working in highly regulated industries such as utilities, financial services, healthcare, insurance, or similar environments.
  • Experience designing maintainable, scalable, and well-documented data infrastructure in cloud-based or modern data platform environments.
  • Experience supporting reproducibility, dataset versioning, release traceability, and audit readiness.
  • Ability to collaborate across multiple technical teams and proactively define expected inputs, outputs, schemas, and interfaces.
  • Strong communication skills with the ability to build trust with stakeholders through clear, accurate, and timely updates.
  • A detail-oriented, governance-minded approach to engineering, with comfort operating in environments that require rigor, documentation, and sign-off discipline.
  • Practical experience with software engineering best practices, including Git-based workflows, code reviews, branching strategies, and release management.

Preferred Qualifications
  • Experience with Palantir Foundry is highly preferred.
  • Experience with GIS technologies and geospatial data platforms.


Additional Information

At Logic20/20, we believe in recognizing and rewarding exceptional talent. Logic20/20 offers a competitive compensation package, with a target base salary range of $156,348 - $175,194 for this role. The final base salary offered is dependent on factors such as relevant experience, skills, qualifications, and location. Eligible employees may also qualify for performance-based bonuses and other incentives.

Similar Jobs

More Jobs at Logic20/20 Inc.

More Information Technology Jobs

Find similar Lead Data Engineer - PySpark/Palantir Foundry jobs: