Job Summary
We are seeking a Data Engineer III to support the development and modernization of a cloud-based enterprise data platform. This role is responsible for designing, developing, and maintaining scalable data pipelines, integrating diverse data sources, and delivering high-quality data products that support analytics and business operations. The ideal candidate will have strong experience with Databricks, Python, PySpark, AWS, Spark-based data processing, and modern data engineering practices within Agile environments.
Key Responsibilities
• Design, develop, and maintain scalable data pipelines to ingest, transform, catalog, and deliver trusted data from multiple enterprise data sources.
• Build and support end-to-end data pipelines for structured, semi-structured, and unstructured data using Apache Spark.
• Develop robust data engineering solutions using Python, PySpark, Databricks, and cloud-native technologies.
• Implement data cataloging, governance, and metadata management processes using enterprise data management tools.
• Build and maintain scalable analytical data stores and modern data lakehouse architectures.
• Monitor data pipelines and implement alerting, automation, and auto-remediation processes to improve reliability and availability.
• Troubleshoot and resolve issues affecting data pipelines, data quality, and analytical platforms.
• Apply security-first principles, automated testing, and data engineering best practices throughout the development lifecycle.
• Collaborate with product managers, data scientists, analysts, and business stakeholders to understand data requirements and deliver scalable solutions.
• Participate in Agile ceremonies and follow SAFe Agile development methodologies.
• Evaluate emerging technologies and recommend improvements to enhance data engineering capabilities and operational efficiency.
• Develop and maintain technical documentation for data pipelines, architectures, and operational processes.
Required Qualifications
• Bachelor's degree in Computer Science, Information Systems, or a related field, or equivalent professional experience.
• 2+ years of experience with Databricks, Collibra, Starburst, or similar enterprise data management platforms.
• 3+ years of experience developing applications using Python and PySpark.
• Experience using Jupyter Notebooks for development, testing, and data analysis.
• Experience working with relational and NoSQL databases, including dimensional modeling and STAR schema design.
• 2+ years of experience with modern data engineering technologies including Amazon S3, Apache Spark, Apache Airflow, lakehouse architectures, real-time databases, Redshift, or Snowflake.
• Experience designing and supporting traditional ETL and Big Data solutions in on-premises or cloud environments.
• Hands-on experience with AWS data engineering services and cloud-native data platforms.
• Experience building end-to-end data pipelines to ingest, process, and transform structured, semi-structured, and unstructured data using Spark architecture.
• Strong analytical, troubleshooting, communication, and collaboration skills.
• Experience working within Agile or SAFe Agile development environments.
Preferred Qualifications
• Experience implementing enterprise data governance and metadata management solutions.
• Experience with cloud-based data lakehouse architectures and distributed data processing frameworks.
• Experience deploying monitoring, alerting, and automated remediation for data platforms.
• Experience collaborating with cross-functional engineering, analytics, and business teams.
• Knowledge of data security, automation, and cloud data engineering best practices.