div:has([data-free-thinking-preview-answer=true])+:is(.text-message,.relative:has(>.text-message))]:-mt-2 grow">
Job Summary
We are seeking an experienced Senior Databricks Data Engineer with strong expertise in designing, developing, and optimizing large-scale data engineering solutions on the Databricks platform. The role will focus on building enterprise-grade data pipelines, processing high-volume datasets, and delivering scalable analytics solutions using PySpark, Spark SQL, Apache Spark, and AWS cloud technologies.
Key Responsibilities
• Design, develop, and maintain scalable data engineering solutions using Databricks, PySpark, Spark SQL, and Apache Spark.
• Build and optimize end-to-end ETL/ELT pipelines integrating Databricks with AWS S3, Snowflake, Airflow, and AWS Glue.
• Process and transform large-scale structured and semi-structured datasets to support enterprise reporting, analytics, and regulatory requirements.
• Develop reusable frameworks and components for data ingestion, cleansing, transformation, validation, reconciliation, enrichment, and business rule implementation.
• Implement robust data quality checks, monitoring, and governance controls to ensure data accuracy and reliability.
• Perform performance tuning and optimization of Spark applications and Databricks workloads to improve scalability and processing efficiency.
• Collaborate with business stakeholders, architects, and analytics teams to understand data requirements and deliver scalable solutions.
• Troubleshoot and resolve production data pipeline issues while maintaining high availability and reliability.
• Support advanced analytics, customer segmentation, campaign analytics, AML and regulatory reporting, and data validation initiatives.
• Contribute to large-scale data migration and modernization programs.
• Support production data platforms and ensure reliable delivery of data engineering solutions.
Required Qualifications
• 7+ years of hands-on experience with Databricks and modern data engineering platforms.
• Strong expertise in PySpark, Spark SQL, Apache Spark, and Databricks Lakehouse architecture.
• Experience building and managing ETL/ELT pipelines on cloud platforms.
• Strong hands-on experience with AWS S3, Snowflake, Apache Airflow, and AWS Glue.
• Strong knowledge of data modeling, data warehousing, and big data processing techniques.
• Experience working with structured and semi-structured data formats, including JSON, Parquet, Avro, and CSV.
• Hands-on experience implementing data quality frameworks and monitoring solutions.
• Strong SQL and data analysis skills.
• Experience developing scalable and production-ready data pipelines.
• Strong troubleshooting and problem-solving skills for production data engineering environments.
Preferred Qualifications
• Experience with regulatory reporting, AML, compliance, or financial data domains.
• Exposure to CI/CD and DevOps practices for data engineering.
• Experience with Delta Lake, Databricks Workflows, and Unity Catalog.
• Knowledge of Agile/Scrum delivery methodologies.
• Experience with customer segmentation analytics.
• Experience with campaign analytics.
• Experience supporting AML and regulatory reporting.
• Experience with enterprise data platforms.
• Experience with data validation and reconciliation.
• Experience with large-scale data migration and modernization programs.
notes:
Mandatory Areas
Must Have Skills
7+ years of hands-on experience with Databricks and modern data engineering platforms.
Strong expertise in: PySpark Spark SQL Apache Spark Databricks Lakehouse Architecture
Experience building and managing ETL/ELT pipelines on cloud platforms.
Strong experience with: AWS S3 Snowflake Apache Airflow AWS Glue Knowledge of data modeling, data warehousing, and big data processing techniques.
Domain Experience (If any) - Good to have healthcare experience
Must have Certifications - None
Location - Remote
Onsite Requirement - Remote
Number of days onsite - Remote