Role Overview
We are seeking a Junior Databricks Engineer to join our team. This role is ideal for a recent master's graduate who is passionate about big data, cloud technologies, and modern data platforms. The successful candidate will work alongside a senior data staff to design, build, and optimize data pipelines and analytics solutions using the Databricks Lakehouse Platform. We're looking for someone who is eager to learn, enjoys solving technical challenges, and thrives in a collaborative, team-oriented environment. This is an excellent opportunity to gain hands-on experience with enterprise-scale data engineering projects while being mentored by experienced professionals.
Key Responsibilities- Develop and maintain data pipelines using Databricks and Apache Spark.
- Assist in building scalable ETL/ELT processes to ingest, transform, and curate data from multiple sources.
- Collaborate with senior engineers to design and implement Lakehouse solutions following Medallion Architecture best practices.
- Write efficient PySpark and SQL code to process large datasets.
- Support data validation, testing, and troubleshooting to ensure data quality and reliability.
- Participate in performance tuning and optimization of Databricks workloads.
- Assist with implementing data governance, security, and access controls.
- Work closely with data analysts, AI engineers, and business stakeholders to understand data requirements.
- Contribute to CI/CD processes, source code management, and deployment automation.
- Participate in Agile ceremonies including sprint planning, daily standups, retrospectives, and code reviews.
- Document technical designs, data flows, and implementation standards.
Required Qualifications- Master's degree in Computer Science, Data Engineering, Data Science, Information Systems, or a related technical discipline.
- Strong understanding of data engineering concepts and relational databases.
- Experience with Python and SQL through academic projects, internships, or personal work.
- Knowledge of Apache Spark fundamentals.
- Familiarity with Databricks or cloud-based analytics platforms.
- Understanding of data modeling and ETL/ELT processes.
- Knowledge of Git and version control best practices.
- Strong analytical, problem-solving, and debugging skills.
- Excellent communication and collaboration abilities.
Preferred Qualifications - Hands-on experience with Databricks through coursework, internships, certifications, or personal projects.
- Experience with Microsoft Azure, AWS, or Google Cloud Platform.
- Familiarity with Delta Lake and Medallion Architecture.
- Knowledge of Azure Data Factory, Microsoft Fabric, or similar data integration tools.
- Exposure to data warehousing concepts and dimensional modeling.
- Experience with Power BI or Tableau.
- Databricks certifications are a plus.
- Understanding of data governance and data quality principles.
- Passionate about modern data engineering technologies.
- Eager to learn from experienced engineers and continuously improve technical skills.
- Strong team player who enjoys collaborating across technical and business teams.
- Curious mindset with a willingness to explore new tools and technologies.
- Highly organized with strong attention to detail.
- Self-motivated, adaptable, and receptive to feedback.
- Committed to delivering high-quality, scalable solutions.
Location: Pittsburgh, PA (relocation assistance is offered)