A Principal Data Engineer is responsible for expanding and optimizing data and data pipeline architecture, as well as optimizing data flow and collection for cross-functional teams. This role involves building and optimizing data systems from the ground up, ensuring that data delivery architecture is consistent throughout ongoing projects.
The Principal Data Engineer supports software developers, database architects, data analysts, and data scientists on data initiatives, ensuring optimal data delivery and architecture. This position requires expertise in data pipeline building and data wrangling, with a focus on enhancing data systems for improved efficiency and performance.
Job Details
Duties & Responsibilities: What major responsibilities does this position have and what percentage of time is spent on completing them? (Typically 5 - 7)
• Develop, construct, test, and maintain data architectures: responsible for creating and maintaining optimal data pipeline architectures necessary for the organization's operational efficiency.
65%
• Collaborate with data scientists and other stakeholders: expected to collaborate with data scientists, business analysts, IT professionals, and management to assist with data-related technical issues and support their data infrastructure needs.
10%
• Assemble and prepare large, complex data sets to meet functional business requirements: required to work with different teams to understand their data needs and ensure those needs are met.
10%
• Build analytics tools: This will involve developing tools that utilize the data pipeline to provide actionable insights into key business performance metrics. 5%
• Maintain data security and compliance: ensure systems meet industry practices and business requirements for data privacy and protection. 5%
• Identify, design, and implement internal process improvements: This includes optimizing data delivery, automating manual processes, and re-designing infrastructure for greater scalability. 3%
• Keep up-to-date with the latest technology trends: Continually learn and adapt to new technologies that could improve existing systems 2%
Knowledge, Skills and Abilities (KSAs): What KSAs are required to perform this job?
• Hands on experience with programming languages (e.g. Python, SQL, GoLang)
• Knowledge with SQL database design concepts and Data Models
• Great numerical and analytical skills
• Degree in Computer Science, IT, or similar field; a Master's is a plus
• Data engineering certification (e.g., Cloud Engineer, Data Engineer) is a plus
• Experience with big data tools: Hadoop, Spark, Kafka, etc.
• Experience with relational SQL and NoSQL databases, including MongoDG, Postgres and/or Cassandra.
• Experience with data pipeline and workflow management tools: Airflow, Dataflow, Composer, etc.
• Experience with GCP, Snowflake, Azure cloud services: BQ, Dataflow, Redshift
• Experience with stream-processing systems: Kafka, Spark-Streaming, etc.
• Good Knowledge of software engineering principles, methodologies and current best practices.
• Able to multitask and shift priorities as needed.
• Able to use BI Tools for data analysis and data QA.
Qualifications
Work Experience &/or Education: What are the minimum education and/or experience requirements necessary to perform this job?
• 9+ years experience and a degree in information technology or computer science with additional vendor-specific certification.
• BS or MS degree in Computer Science or a related technical field
• 4+ years of Python or Java development experience
• 4+ years of SQL experience (No-SQL experience is a plus)
• 4+ years of experience with schema design and dimensional data modeling
• Ability in managing and communicating data warehouse plans to internal clients
• Experience designing, building, and maintaining data processing systems
• Experience working with a cloud platform such as Snowflake / Azure or Databricks