Full Job Description
As a Lead Data Engineer - Python/PySpark/Databricks/AWS/AI at JPMorganChase within the Consumer & Community Banking, you are an integral part of an agile team that works to enhance, build, and deliver data collection, storage, access, and analytics solutions in a secure, stable, and scalable way. As a core technical contributor, you are responsible for maintaining critical data pipelines and architectures across multiple technical areas within various business functions in support of the firm’s business objectives.
Job responsibilities
• Generates data models for their team using firmwide tooling, linear algebra, statistics, and geometrical algorithms
• Delivers data collection, storage, access, and analytics data platform solutions in a secure, stable, and scalable way
• Implements database back-up, recovery, and archiving strategy
• Evaluates and reports on access control processes to determine effectiveness of data asset security with minimal supervision
• Uses enterprise-authorized AI capabilities within the work environment to accelerate data platform and model design analysis and documentation, validating outputs and handling data according to sensitivity and security requirements.
• Applies reuse-first, AI-assisted practices within delivery and operational routines (e.g., backup/recovery validation and access control review support), ensuring traceability/auditability and alignment to resiliency and security expectations.
Required qualifications, capabilities, and skills
• Formal training or certification on Data Science engineering concepts and 5+ years applied experience
• Expertise with Python, PySpark, Databricks, Snowflake, AWS and AI
• Working experience with both relational and NoSQL databases
• Experience and proficiency across the data lifecycle
• Experience with database back-up, recovery, and archiving strategy
• Proficient knowledge of linear algebra, statistics, and geometrical algorithms
• Demonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity.
• Ability to review and validate AI-assisted outputs (e.g., model/design summaries or operational checklists) before use, escalating when uncertain and following data handling requirements.
Preferred qualifications, capabilities, and skills
• Exposure to cloud technologies
• Hands-on experience on Kafka or any streaming technology
• Hands-on experience in Splunk, Dynatrace tools
• Exposure to AI Driven development