Role Overview:Responsible for developing, deploying, and optimizing scalable data pipelines, data models, and analytics solutions to support enterprise-wide business insights and decision-making. This role requires expertise in cloud data engineering, data warehousing, advanced SQL, and Python/PySpark transformations, with the ability to translate data into actionable business recommendations. The ideal candidate brings deep technical experience, strong analytical skills, and cross-functional collaboration abilities, preferably within the pharmaceutical or healthcare domain.
Key Responsibilities:- Develop and maintain scalable ETL/ELT pipelines using Python, PySpark, and SQL for structured and unstructured data ingestion.
- Build robust data infrastructure on cloud platforms such as Azure Fabric, Synapse, Databricks, and AWS.
- Implement end-to-end automation for data ingestion, streaming, scheduling, and monitoring across Azure Functions, Data Factory, ADLS Gen2, and Power BI.
- Optimize large-scale big data architectures for performance and scalability.
- Design and maintain multidimensional models (Star/Snowflake Schema) including fact and dimension tables, views, and stored procedures.
- Manage complex datasets ensuring alignment with functional and non-functional requirements.
- Learn foundational SAP S4/HANA and BW4 data models to support financial analytics.
- Utilize GitHub and Azure DevOps for CI/CD, version control, and automated deployments.
- Manage artifact deployment across environments with risk assessments and impact analysis.
- Implement process improvements and workflow automation.
- Partner with business stakeholders to gather requirements and translate them into analytical solutions.
- Develop analytics and reporting solutions supporting Global Finance transformation strategy.
- Develop dashboards and BI tools using Power BI/Tableau, translating outcomes into actionable insights with KPIs.
- Troubleshoot data-related issues and perform root cause analysis.
Required Skills:- 5-8 years of hands-on experience in data engineering and cloud data warehousing (Azure Synapse, Fabric, Databricks, Redshift, or Snowflake).
- Expert SQL proficiency across relational databases and cloud data warehouses, plus strong Python and PySpark skills for data transformation and pipeline development.
- Data modelling expertise including dimensional modelling (facts, dims), normalization/de-normalization, views, and stored procedures for automation.
- End-to-end ownership of project delivery with proven ability to translate business problems into scalable analytical solutions.
- Big data architecture experience building and optimizing large-scale pipelines across structured and unstructured datasets.
- Visualization expertise with Power BI and/or Tableau to present complex findings clearly.
- Strong documentation and communication skills to support operational standards and collaborate with cross-functional teams.
- Experience with root cause analysis, continuous improvement, and Agile delivery methodologies.
Qualifications:- Bachelor's or Master's degree in Technology, Computer Science, Engineering, or related field.
Preferred Skills:- Experience with SQL and NoSQL databases (AWS Redshift, Postgres, Databricks etc.).
- Pharmaceutical or healthcare dataset experience.
- Cloud expertise in AWS/Azure; Microsoft Fabric experience is a strong advantage.
- CI/CD experience using GitHub or Azure DevOps.
- Proficiency in Python/PySpark, R, Scala, or additional scripting languages.
- AI & Gen AI - Products & Tools.