Job Summary
The Data Engineer will design, develop, and support next-generation, cloud-native data platforms that enable real-time and batch data movement at enterprise scale. The role will focus on building streaming and batch frameworks using Apache Spark, Apache Kafka, Apache Flink, and Airflow, along with cloud-native solutions on Microsoft Azure and Kubernetes. The engineer will contribute to platform modernization, automation, cloud migration, performance optimization, and operational excellence across distributed data platforms.
Key Responsibilities
• Design, develop, and support scalable real-time and batch data pipelines using Apache Spark, Apache Flink, Apache Kafka, and Airflow.
• Build and enhance metadata-driven, self-service data integration platforms and reusable connectors.
• Develop and maintain source and target connectors for relational databases, file systems, Kafka, Cassandra, YugabyteDB, and other enterprise data stores.
• Design, deploy, and operate Kubernetes-based streaming and batch processing platforms.
• Lead cloud migration initiatives from on-premises environments to Microsoft Azure.
• Drive performance tuning, scalability optimization, reliability improvements, and operational excellence across large-scale data workloads.
• Develop cloud-native solutions supporting enterprise data movement and processing.
• Collaborate with product owners, architects, cloud engineering teams, and business stakeholders to deliver enterprise-scale data solutions.
• Contribute to platform modernization, automation, CI/CD implementation, and engineering best practices.
• Support troubleshooting, production stability, and continuous improvement initiatives for critical data platforms.
Required Qualifications
• 8+ years of overall Software Engineering or Data Engineering experience.
• 5+ years of hands-on experience with Apache Spark for large-scale batch and streaming data processing.
• 5+ years of hands-on experience with Apache Kafka, including event-driven architectures and real-time data streaming solutions.
• 3+ years of hands-on experience with Apache Flink for stream processing and real-time analytics workloads.
• 5+ years of experience developing applications using Java and/or Scala.
• 5+ years of experience writing complex SQL queries and optimizing database performance.
• 3+ years of experience with Kubernetes and containerized application deployment.
• 3+ years of experience designing and implementing solutions on Microsoft Azure Cloud.
• 3+ years of experience building and supporting distributed data platforms using Cassandra, YugabyteDB, PostgreSQL, or similar databases.
• 3+ years of experience building event-driven architectures and streaming applications.
• 2+ years of experience implementing CI/CD pipelines, source control, and deployment automation using GitLab, GitHub, or similar tools.
• 2+ years of experience with cloud-native deployment patterns, containerization, and platform automation.
• Demonstrated experience in performance tuning, troubleshooting, and supporting mission-critical data platforms.
• Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent work experience.
Preferred Qualifications
• Experience with Apache Airflow orchestration.
• Experience with metadata-driven data platforms and self-service ingestion frameworks.
• Experience with enterprise cloud migration and modernization programs.
• Experience with platform engineering and internal developer platforms.
• Experience leading technical initiatives, mentoring engineers, or serving as a technical lead.
• Knowledge of DevSecOps, Infrastructure as Code, and observability frameworks.