Job DescriptionAs a Lead Data Engineer joining our AI & Analytics practice, you'll be responsible for delivering client value and ensuring high client satisfaction. You'll be expected to be adept at recognizing, subscribing, and applying best practices, methodologies, tools, and techniques to meet client requirements, timelines, and budgets.
The right candidate will be highly hands-on, comfortable working across cross-functional teams, and motivated by clean architecture, disciplined engineering practices, and well-governed delivery processes.
Role Responsibilities- Design, enhance, and maintain production-grade data pipelines that support model output aggregation and downstream risk analysis.
- Establish and mature repository governance practices, including branching strategy, pull request standards, merge policies, release tagging, and version control workflows.
- Own release engineering practices that support reproducible, traceable, and auditable production releases.
- Refactor and improve existing pipeline code to increase modularity, maintainability, scalability, and documentation quality.
- Develop configuration-driven pipeline patterns that create consistency and reduce ambiguity across environments and releases.
- Support testing, validation, benchmarking, and change management practices for critical data pipelines.
- Partner with data scientists, machine learning engineers, data engineers, product stakeholders, and other technical teams to align on schemas, interfaces, inputs, and delivery expectations.
- Translate complex technical concepts into clear updates for both technical and non-technical stakeholders.
- Help bring structure to existing codebases, repositories, and engineering workflows that may need stronger governance or standardization.
- Contribute to engineering best practices across a highly regulated, audit-sensitive delivery environment.
Qualifications- 10-15+ years of data engineering, data science, machine learning engineering, and/or relevant experience using Python.
- Experience leading technical teams and overseeing enterprise-scale data initiatives.
- Strong expertise in PySpark, SQL, and cloud services.
- Demonstrated ability to improve, refactor, or stabilize existing codebases and pipeline environments.
- Experience with cloud-optimized datasets, efficient partitioning strategies, and large-scale spatial operations.
- Strong understanding of how machine learning model outputs flow into downstream data pipelines, platforms, or production systems.
- Experience working in highly regulated industries such as utilities, financial services, healthcare, insurance, or similar environments.
- Experience designing maintainable, scalable, and well-documented data infrastructure in cloud-based or modern data platform environments.
- Experience supporting reproducibility, dataset versioning, release traceability, and audit readiness.
- Ability to collaborate across multiple technical teams and proactively define expected inputs, outputs, schemas, and interfaces.
- Strong communication skills with the ability to build trust with stakeholders through clear, accurate, and timely updates.
- A detail-oriented, governance-minded approach to engineering, with comfort operating in environments that require rigor, documentation, and sign-off discipline.
- Practical experience with software engineering best practices, including Git-based workflows, code reviews, branching strategies, and release management.
Preferred Qualifications- Experience with Palantir Foundry is highly preferred.
- Experience with GIS technologies and geospatial data platforms.
Additional InformationAt Logic20/20, we believe in recognizing and rewarding exceptional talent. Logic20/20 offers a competitive compensation package, with a target base salary range of $156,348 - $175,194 for this role. The final base salary offered is dependent on factors such as relevant experience, skills, qualifications, and location. Eligible employees may also qualify for performance-based bonuses and other incentives.