Bachelor's degree in computer science, Information Systems, Engineering, Data Analytics, or a related field, or equivalent experience.
5+ years of data engineering experience with Azure, Databricks, or similar platforms, including 2+ years as a senior engineer or mentor.
Expertise in Databricks, Spark, PySpark, Delta Lake, Unity Catalog, SQL, Python, and Lakehouse architecture.
Knowledge of Azure Data Lake Storage, Azure SQL, Key Vault, ETL/ELT processes, dimensional modeling, and cloud security fundamentals.
Experience designing data solutions for analytics, machine learning, AI, and secure data use cases, including feature engineering and model-ready data preparation.
Databricks Certified Data Engineer Associate certification, with a preference for Professional certification.
Responsibilities
Design and support secure cloud-based data solutions for analytics and AI initiatives.
Lead the team of data engineers and oversee work planning, design reviews, and technical guidance.
Establish engineering standards and best practices for data ingestion, transformation, and governance.
Optimize production data pipelines, ensuring performance and reliability across Databricks and Azure.
Drive architectural decisions and design AI-ready data assets for enterprise use.
Champion data governance through implementation of Unity Catalog and access controls.
Collaborate with cross-functional teams to deliver trusted enterprise data solutions.
Benefits
Professional development opportunities and mentorship.
Flexible working arrangements and remote work options.
Access to cutting-edge technology and tools.
Collaborative and innovative team environment.
Health and wellness programs and resources.
Full Job Description
EDUCATION AND EXPERIENCE
Bachelor's degree in computer science, Information Systems, Engineering, Data Analytics, or a related field, or an equivalent combination of education and experience.
Five (5) or more years of hands-on data engineering experience using Azure, Databricks, or comparable cloud data platforms, including two (2) or more years serving as a senior engineer, technical lead, or mentor on data engineering teams or projects.
Demonstrated experience in Databricks, Spark, PySpark, Delta Lake, Unity Catalog, SQL, Python, notebooks, jobs, clusters, and Lakehouse architecture is required.
Experience with Azure Data Lake Storage (ADLS), Azure SQL, Key Vault, managed identities, ETL/ELT processes, dimensional modeling, and cloud security fundamentals is required.
Experience designing data solutions that support analytics, machine learning, artificial intelligence, governed enterprise reporting, and secure data consumption use cases is also required, including knowledge of AI/ML data preparation practices such as feature engineering, model-ready data design, data quality, lineage, and secure access to sensitive data.
Databricks Certified Data Engineer Associate certification or an equivalent current Databricks data engineering certification is required;
Databricks Certified Data Engineer Professional certification is preferred.
JOB SUMMARY
The Senior Databricks Data Engineer designs, develops, and supports scalable, secure, and governed cloud-based data solutions that enable enterprise analytics, business intelligence, regulatory reporting, artificial intelligence initiatives, and data products. This role serves as a senior technical resource for the enterprise data platform and provides day-to-day technical leadership to a team of data engineers, including work planning and assignment, design and code review, mentoring, and escalation support for complex production issues.
In addition to building, maintaining, and optimizing production data pipelines, this position leads the design and development of the organization's enterprise data platform using Azure and Databricks. The position applies and establishes modern Lakehouse engineering practices and technologies, including Databricks, Apache Spark, PySpark, Delta Lake, Unity Catalog, Azure Data Lake Storage, SQL, and Python, with an emphasis on data governance, performance, automation, security, and production reliability. The Senior Databricks Data Engineer also drives data solution architecture by establishing and applying Lakehouse and Medallion architecture patterns, developing governed and AI-ready data assets, creating reusable engineering frameworks, and supporting secure and consistent data consumption across the organization.
ESSENTIAL FUNCTIONS
Provide day-to-day technical leadership and direction to a team of data engineers, including assigning and prioritizing work, estimating effort, tracking delivery, and removing technical obstacles in partnership with the manager.
Mentor and develop engineers through design guidance, pair programming, knowledge sharing, and structured onboarding of new team members, contractors, and partners.
Lead technical design sessions and establish engineering standards, patterns, and best practices for ingestion, transformation, governance, deployment, observability, and data consumption.
Own technical quality across the team by conducting design and code reviews and enforcing testing, documentation, and release standards.
Partner with data, analytics, application, infrastructure, security, and AI teams to architect the overall enterprise data solution using Azure and Databricks.
Design and maintain batch and near-real-time pipelines using Databricks, ADF, Spark/PySpark, SQL, REST APIs, and ADLS.
Build governed Bronze, Silver, and Gold data layers using Delta Lake, Medallion Architecture, schema enforcement, incremental processing, and data quality controls.
Implement Unity Catalog, RBAC, lineage, classification, and secure data access/sharing across Databricks environments.
Deliver reusable data products supporting Power BI, enterprise and regulatory reporting, analytics, and AI/ML.
Design AI-ready data pipelines and curated datasets that support machine learning, generative AI, feature engineering, retrieval-augmented generation, vector search, and advanced analytics use cases.
Support Databricks AI and ML capabilities, including MLflow, feature engineering patterns, model-ready datasets, model lifecycle support, model monitoring, and responsible AI governance controls.
Manage orchestration, scheduling, monitoring, alerting, error handling, testing, and production support, and serve as a senior escalation point for complex production incidents and root cause analysis.
Optimize Spark, SQL, clusters, storage, and compute for performance, reliability, and cost efficiency.
Implement solutions using Azure DevOps, Git, CI/CD, automated testing, and environment promotion.
Collaborate with analytics, application, security, infrastructure, and business teams to deliver trusted enterprise data.
Contribute to reference architectures, reusable frameworks, platform standards, and implementation patterns for ingestion, transformation, governance, deployment, observability, and data consumption.
Support platform planning activities by providing roadmap input, level-of-effort estimates, resource considerations, and assessments of technical risk.
REQUIRED SKILLS
Ability to lead, mentor, and develop technical staff, and to guide the work of other engineers without direct supervisory authority.
Ability to lead technical design discussions, evaluate alternatives, build consensus, and make sound architecture and implementation decisions.
Ability to plan, prioritize, and coordinate work across multiple engineers, concurrent projects, and competing deadlines.
Ability to communicate technical concepts clearly and collaborate with data analysts, BI developers, data scientists, IT teams, business stakeholders, and other technical partners.
Advanced proficiency with Databricks, Delta Lake, Unity Catalog, Medallion architecture, and Lakehouse design principles.
Advanced skills developing, troubleshooting, and optimizing Spark and PySpark workloads.
Strong programming and query-development skills using SQL and Python for data engineering, transformation, and automation.
Hands-on knowledge of data pipelines in Databricks, Azure Data Lake Storage (ADLS), Azure SQL, Key Vault, and managed identities.
Ability to design, build, maintain, and optimize scalable ETL/ELT pipelines and production data workflows.
Knowledge of dimensional modeling, data integration, data layers, and reusable enterprise data architecture patterns.
Understanding of data governance, lineage, access controls, data quality, sensitive data protection, and cloud security practices.
Ability to identify and resolve performance issues involving Spark workloads, SQL queries, clusters, and data pipelines.
Working knowledge of Azure DevOps, Git, source control, automated testing, CI/CD pipelines, release management, and environment promotion.
Knowledge of preparing governed, high-quality, model-ready datasets for machine learning and AI applications.
Ability to diagnose data pipeline failures, identify root causes, implement solutions, and maintain reliable production environments.
Ability to analyze complex technical and data issues and develop scalable, practical solutions.