Job DescriptionWe're looking for a Data Engineer to join our Databricks practice and help design, build, and optimize Lakehouse-based data platforms for clients across industries - from public sector agencies to healthcare, financial services, and manufacturing. You'll work hands-on with the Databricks Data Intelligence Platform to turn messy, disconnected client data into governed, trustworthy, analytics- and AI-ready assets.
This is a client-facing consulting role. You'll partner with solution architects, data scientists, and project leads to gather requirements, design pipelines, and deliver production-grade solutions - then explain what you built and why it matters in language business stakeholders actually understand.
What You'll Do - Design, build, and optimize scalable ETL/ELT pipelines on Databricks using PySpark, Spark SQL, and Delta Lake
- Implement Medallion (Bronze/Silver/Gold) architecture patterns, applying data quality checks, schema evolution, and enforcement along the way
- Build declarative pipelines with Delta Live Tables (DLT) and ingest streaming/incremental data using Auto Loader and Structured Streaming
- Orchestrate and monitor production workloads using Databricks Workflows, integrating with tools like Airflow or Azure Data Factory where needed
- Configure and maintain Unity Catalog for data governance - catalogs, schemas, access controls, lineage, and PII masking
- Partner with data scientists to prepare feature-engineered, ML-ready datasets and support model deployment workflows using MLflow
- Tune cluster configuration, job design, and Photon/serverless compute for performance and cost efficiency
- Build and maintain CI/CD pipelines for Databricks notebooks, jobs, and asset bundles (Git-based workflows, Azure DevOps, GitHub Actions, or similar)
- Query, profile, and assess the quality of large, complex datasets from a wide variety of source systems
- Collaborate with solution leads, architects, and project managers on solution design and technical architecture decisions
- Participate directly in client-facing work: requirements gathering, solution reviews, and translating technical tradeoffs into plain-language business impact
- Document solutions clearly - architecture diagrams, data flow documentation, code comments, and runbooks
QualificationsRequired- Bachelor's degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience)
- 2+ years of hands-on data engineering experience, including production work on the Databricks platform
- Strong hands-on experience with PySpark and Spark SQL
- Practical experience with Delta Lake fundamentals - ACID transactions, OPTIMIZE/Z-ORDER, partitioning, and schema evolution
- Solid SQL skills across relational platforms (SQL Server, Postgres, Oracle, Snowflake, etc.)
- Experience with at least one major cloud platform (Azure, AWS, or GCP) and its data services
- Working knowledge of data modeling (dimensional modeling, 3NF) and ETL/ELT design principles
- Strong communication skills and comfort working directly with clients and non-technical stakeholders
- A collaborative, detail-oriented mindset with a bias toward solution quality and follow-through
Preferred / Nice-to-Have - Databricks Certified Data Engineer Associate or Professional
- Experience with Unity Catalog, Delta Live Tables, and Auto Loader in production environments
- Exposure to MLflow, Feature Store, or Databricks Vector Search for AI/ML-enabled use cases
- Experience with Databricks Asset Bundles and CI/CD tooling (GitHub Actions, Azure DevOps, GitLab)
- Familiarity with Terraform or other infrastructure-as-code tooling
- Experience with dbt, Kafka/Event Hubs, or BI tools (Power BI, Tableau) connected to Databricks
- Docker/Kubernetes experience for containerized workloads
- Prior consulting experience, or comfort moving across multiple client engagements and industries