Job Summary
Senior Databricks Architect responsible for leading the architecture and technical execution of large-scale data modernization and migration initiatives within the Retail/Grocery domain. The role focuses on designing target-state Databricks Lakehouse architecture, defining migration strategies, guiding data engineering teams, and modernizing Talend-based ETL/ELT workloads onto Databricks. The position requires strong hands-on expertise in Databricks, Apache Spark, PySpark, SQL, Delta Lake, cloud data platforms, data architecture, ETL modernization, and migration strategy, along with a strong understanding of Retail/Grocery data domains.
Key Responsibilities
• Lead the end-to-end architecture for migrating Talend ETL/ELT workloads to Databricks.
• Define target-state Databricks Lakehouse architecture, including data ingestion, transformation, storage, orchestration, governance, and consumption layers.
• Design scalable, secure, highly available, and cost-optimized data solutions.
• Define architecture standards, design patterns, coding standards, and best practices for Databricks development.
• Establish reusable migration frameworks and patterns for converting Talend jobs into Databricks/Spark-based pipelines.
• Evaluate existing Talend jobs and determine appropriate migration strategies, including re-platforming, re-engineering, consolidation, or retirement.
• Analyze Talend workflows, jobs, mappings, dependencies, schedules, and data transformations.
• Develop migration strategies for batch and incremental data processing workloads.
• Translate Talend transformations and business rules into PySpark, SQL, and Databricks implementations.
• Identify opportunities to simplify and optimize legacy ETL processes during migration.
• Define data reconciliation and validation strategies to ensure parity between Talend and Databricks outputs.
• Establish migration sequencing based on business criticality, dependencies, complexity, and risk.
• Support migration of high-volume and business-critical data pipelines.
• Design and implement solutions using Databricks, Apache Spark, Delta Lake, PySpark, and SQL.
• Design Medallion Architecture using Bronze, Silver, and Gold layers.
• Implement incremental processing, CDC, SCD Type 1/Type 2, data quality, error handling, and audit frameworks.
• Optimize Spark workloads, Delta tables, SQL queries, partitioning, clustering, and job execution.
• Leverage Databricks Workflows/Jobs, Unity Catalog, and modern Databricks data engineering capabilities.
• Design data pipelines integrating structured and semi-structured data from databases, files, APIs, and enterprise applications.
• Establish monitoring, logging, alerting, operational support, and performance management practices.
• Design integrations between Databricks and cloud storage, databases, messaging platforms, APIs, and enterprise applications.
• Define secure connectivity, networking, secrets management, and access-control patterns.
• Work with business and technology stakeholders to understand Retail/Grocery data requirements and translate them into scalable data architecture.
• Provide technical leadership and mentoring to data engineers and developers.
• Conduct architecture, design, and code reviews.
• Troubleshoot complex production and performance issues.
• Collaborate with onshore/offshore teams, enterprise architects, infrastructure teams, security teams, and business stakeholders.
• Prepare architecture diagrams, technical design documents, migration plans, and implementation standards.
• Communicate technical risks, dependencies, and recommendations to client leadership.
Required Qualifications
• 10+ years of overall IT experience.
• 5+ years of experience in Databricks, cloud data engineering, and/or data architecture.
• Strong hands-on experience with Databricks.
• Strong experience with Apache Spark and PySpark.
• Strong experience with Delta Lake.
• Experience with Databricks Workflows/Jobs and Databricks SQL.
• Experience with Unity Catalog.
• Strong experience with performance tuning and optimization.
• Strong understanding of Medallion and Lakehouse architectures.
• Strong SQL and data modeling skills.
• Strong experience with ETL/ELT architecture.
• Experience with batch and incremental data processing.
• Experience implementing CDC and SCD Type 1/Type 2.
• Experience with data quality and reconciliation frameworks.
• Experience processing large-volume datasets.
• Experience leading complex data migration or modernization initiatives.
• Strong understanding of data architecture, integration, governance, and security principles.
• Experience with at least one major cloud platform, preferably Azure or AWS.
• Strong Retail/Grocery domain understanding or experience working with retail data environments.
Preferred Qualifications
• Experience with ADLS or S3.
• Experience with Azure Data Factory.
• Experience with Kafka.
• Experience with REST APIs and cloud-native services.
• Experience integrating Databricks with enterprise applications and relational databases.
• Experience with Retail/Grocery data domains such as Product/Item, Store/Location, Customer/Loyalty, Sales/Transactions, Pricing, Promotions, Inventory, Supply Chain, Vendors/Suppliers, Orders, Forecasting, and Merchandising.
• Experience with high-volume transaction processing, store-level data, product hierarchies, promotions, pricing, inventory movements, and customer analytics.