Job Summary
We are seeking a Data Engineer with expertise in Azure Databricks to support the modernization of enterprise data platforms. The ideal candidate will have experience migrating Oracle PL/SQL-based ETL processes to Databricks, developing scalable cloud-based data pipelines, and ensuring data quality, reliability, and performance.
Key Responsibilities
• Gather business and technical requirements from stakeholders.
• Analyze Oracle SQL and PL/SQL ETL processes for migration.
• Migrate Oracle ETL workflows to Azure Databricks using Spark SQL and PySpark.
• Design, develop, test, deploy, and maintain scalable data pipelines.
• Create technical design documents and data mappings.
• Validate migrated data with business users and SMEs.
• Implement incremental data processing using Delta Lake.
• Convert complex Oracle PL/SQL logic into Databricks pipelines.
• Perform data validation and reconciliation between source and target systems.
• Monitor, troubleshoot, and enhance production ETL pipelines.
• Optimize pipeline performance and resolve production issues.
• Participate in source control, CI/CD, and release management activities.
• Manage assigned projects and deliver high-quality solutions.
Required Qualifications
• Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field.
• 4-6 years of experience in ETL, database development, or data engineering.
• 4+ years of advanced SQL development and query optimization.
• 3+ years of Oracle PL/SQL development experience.
• 3+ years of Azure Databricks development experience.
• Experience migrating legacy ETL processes to cloud-based data platforms.
• Strong experience with Spark SQL, PySpark, Delta Lake, and Azure Data Lake Storage (ADLS).
• Experience with Git, source control, and CI/CD processes.
• Strong analytical and problem-solving skills.
• Excellent communication and collaboration skills.
Preferred Qualifications
• Experience in healthcare, biomedical research, or clinical data environments.
• Experience converting Oracle PL/SQL packages, procedures, and functions into Databricks pipelines.
• Experience with Azure Data Factory or Databricks Workflows.
• Knowledge of dimensional data modeling and enterprise data warehousing.
• Experience with Python.
• Experience with Power BI or Tableau.
• Familiarity with AI or machine learning-enabled data engineering solutions.