Specialist - Data Engineering

LTM

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in data engineering or related fields
  • Proficient in big data frameworks, particularly Apache Spark and PySpark
  • Strong programming capabilities in Python, with focus on data processing
  • Expertise in SQL and data schema management, especially HiveQL
  • Experience in cloud data infrastructure using AWS, Azure, or GCP
  • Solid understanding of data warehousing and dimensional modeling principles
  • Familiarity with CI/CD practices and automation tools for deployment

Responsibilities

  • Design, build, and maintain scalable ETL/ELT data pipelines using PySpark
  • Deploy and manage cloud data infrastructure components in AWS, Azure, or GCP
  • Optimize data storage and accessibility in Apache Hive and cloud data lakes
  • Proactively identify and resolve Spark job performance issues
  • Develop solutions for high-volume data ingestion from diverse sources
  • Implement automated workflows for reliable data delivery using tools like Apache Airflow
  • Collaborate with data scientists and analysts to create efficient data solutions

Benefits

  • Access to advanced tools and technologies in a dynamic work environment
  • Opportunities for professional development and training
  • Collaboration with cross-functional teams and experts in data science
  • Flexible work arrangements and emphasis on work-life balance
  • Possibility of gaining valuable cloud certifications relevant to career growth
Full Job Description
Role description

Karat Interview Process

Job Summary

We are seeking a highly skilled and motivated Data Engineer to play a pivotal role in designing building and optimizing our nextgeneration scalable data pipelines This position requires expertise in processing massive datasets using cuttingedge technologies like Apache Spark PySpark and Hive within a dynamic cloud environment Your primary objective will be to ensure the utmost data reliability speed and efficiency providing a robust foundation for downstream business intelligence and advanced analytics initiatives

Key Responsibilities

Data Pipeline Development Maintenance Design build and maintain highly scalable and efficient ETLELT data pipelines utilizing PySpark and Spark SQL for complex data transformations

Cloud Data Infrastructure Management Deploy manage and scale critical data infrastructure components on leading cloud platforms such as Amazon Web Services AWS eg EMR Glue Microsoft Azure eg Databricks Synapse or Google Cloud Platform GCP

Data Warehousing Storage Optimization Strategically manage data layout partitioning and indexing within Apache Hive and various cloud data lake solutions to optimize performance and accessibility

Performance Tuning Optimization Proactively identify and resolve performance bottlenecks in Spark jobs leveraging Spark UI for indepth analysis effectively managing data skewness and optimizing memory utilization

Diverse Data Integration Develop robust solutions for ingesting highvolume and diverse datasets from both structured relational databases and unstructured flat files into our data ecosystem

Automated Workflow Orchestration Implement and manage automated data workflows using industrystandard scheduling tools like Apache Airflow or platformnative schedulers ensuring timely and reliable data delivery

Strategic Collaboration Partner closely with data scientists business analysts and crossfunctional enterprise teams to translate complex business requirements into technically sound and efficient data solutions

Required Core Technical Skills

Big Data Frameworks Expertise Demonstrated high proficiency in Apache Spark architecture including a deep understanding of drivers executors and Directed Acyclic Graphs DAGs

Advanced Programming Exceptional coding skills in Python and extensive experience with the PySpark API for developing intricate data transformations and processing logic

Querying Schema Management Strong command of HiveQL and ANSI SQL coupled with expertise in data partitioning techniques and effective schema definition

Optimized Storage Formats Indepth understanding and practical experience with optimized big data storage file formats such as Parquet ORC and Avro

Cloud Ecosystem Development Handson development experience utilizing cloudnative big data utilities eg AWS EMR Azure Databricks within major cloud platforms

Data Warehousing Fundamentals Solid foundation in Dimensional Data Modeling including Star and Snowflake schemas and practical experience with Data Lakes concepts and implementation

Preferred Qualifications

CICD DevOps Automation Experience with Continuous IntegrationContinuous Deployment CICD practices and automation tools like Git Jenkins or Ansible

NoSQL Database Integration Exposure to and experience with NoSQL databases such as HBase Cassandra or MongoDB

Professional Cloud Certifications Relevant professional cloud certifications eg AWS Certified Data Engineer Microsoft Certified Azure Data Engineer Associate are highly valued

Similar Jobs

More Jobs at LTM

More Information Technology Jobs

Find similar Specialist - Data Engineering jobs: