Job Description:
Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java)Join a team building next-generation, cloud-native data platforms powering real-time and batch data movement at enterprise scale. You will design and develop streaming and batch frameworks, Kafka/Flink/Spark-based processing engines, and cloud-native solutions on Azure and Kubernetes. Ideal candidates have strong distributed systems experience and a passion for platform engineering, automation, and data modernization.
Job Description
Job Title
Data Engineer - Data Platform (Spark/Kafka/Flink/Scala/Java)
Day to Day Job Duties:- Design, develop, and support scalable real-time and batch data pipelines using Apache Spark, Apache Flink, Apache Kafka, and Airflow.
- Build and enhance metadata-driven self-service data integration platforms and reusable connectors.
- Develop and maintain source and target connectors for relational databases, file systems, Kafka, Cassandra, YugabyteDB, and other enterprise data stores.
- Design, deploy, and operate Kubernetes-based streaming and batch processing platforms.
- Lead cloud migration initiatives from on-premise environments to Microsoft Azure.
- Drive performance tuning, scalability optimization, reliability improvements, and operational excellence across large-scale data workloads.
- Develop cloud-native solutions supporting enterprise data movement and processing.
- Collaborate with product owners, architects, cloud engineering teams, and business stakeholders to deliver enterprise-scale data solutions.
- Contribute to platform modernization, automation, CI/CD implementation, and engineering best practices.
- Support troubleshooting, production stability, and continuous improvement initiatives for critical data platforms.
Basic Qualifications:- (What are the skills required for this job with minimum years of experience on each)
- Minimum 8+ years of overall Software Engineering or Data Engineering experience.
- Minimum 5+ years of hands-on experience with Apache Spark for large-scale batch and streaming data processing.
- Minimum 5+ years of hands-on experience with Apache Kafka including event-driven architectures and real-time data streaming solutions.
- Minimum 3+ years of hands-on experience with Apache Flink for stream processing and real-time analytics workloads.
- Minimum 5+ years of experience developing applications using Java and/or Scala.
- Minimum 5+ years of experience writing complex SQL queries and optimizing database performance.
- Minimum 3+ years of experience with Kubernetes and containerized application deployment.
- Minimum 3+ years of experience designing and implementing solutions on Microsoft Azure Cloud.
- Minimum 3+ years of experience building and supporting distributed data platforms using Cassandra, YugabyteDB, PostgreSQL, or similar databases.
- Minimum 3+ years of experience building event-driven architectures and streaming applications.
- Minimum 2+ years of experience implementing CI/CD pipelines, source control, and deployment automation using GitLab, GitHub, or similar tools.
- Minimum 2+ years of experience with cloud-native deployment patterns, containerization, and platform automation.
- Demonstrated experience in performance tuning, troubleshooting, and supporting mission-critical data platforms.
Travel:- Minimal travel required. Travel may be necessary based on project and stakeholder requirements.
Degree:- Bachelor's degree in Computer Science, Information Technology, Engineering, or equivalent work experience.
Nice to Have (But Not a Must)- Experience with Apache Airflow orchestration.
- Experience with metadata-driven data platforms and self-service ingestion frameworks.
- Experience with enterprise cloud migration and modernization programs.
- Experience with platform engineering and internal developer platforms.
- Experience leading technical initiatives, mentoring engineers, or serving as a technical lead.
- Knowledge of DevSecOps, infrastructure as code, and observability frameworks.