Role description
Expertise in Designing and developing scalable Apache spark ETL based Data processing pipelines
Strong commandline knowledge in UnixLinux with Shell scripting using Bash Kornshell or Perl and File processing using awk scripts
Expertise in SQL querying and complex joins
Implementing comprehensive Spark based Data validation frameworks transforming large volumes of Financial data within the Project lifecycle
Expertise with complex Data workflows with Apache AirFlow managing task dependencies SLAs etc to ensure timely data delivery and corresponding automated validation controls
Strong Analytical skills and expertise on SparkSQL for Data analysis and validation ensuring the delivery of clean queryready datasets for business consumption
Expertise in Data quality checks and monitoring
Karat interview Process
Reference Sabin