Java Spark Engineer

Compunnel

$125K — $150K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or related field
  • 7+ years of professional Java development experience
  • 5+ years of hands-on Apache Spark experience in production environments
  • Expert understanding of distributed systems and data processing
  • Strong SQL skills and familiarity with modern data storage formats
  • Experience with cluster managers like YARN and Kubernetes
  • Proficiency with Kafka and CI/CD practices

Responsibilities

  • Architect and build scalable, fault-tolerant data pipelines using Apache Spark and Java
  • Lead design and implementation of batch and streaming ETL/ELT systems
  • Perform deep performance tuning and job cost reduction
  • Establish coding standards and lead code reviews
  • Drive technical decisions for data architecture and pipeline orchestration
  • Mentor junior engineers and serve as a technical escalation point
  • Partner with teams to translate requirements into scalable data systems
  • Own production reliability and incident response

Benefits

  • Collaborative work environment with strong mentorship opportunities
  • Impactful role in shaping data architecture and driving technical decisions
  • Opportunity to work with cutting-edge technologies and frameworks
  • Possibility to contribute to capacity planning and cost optimization
  • Engagement with interdisciplinary teams to enhance data system scalability
Full Job Description
Job Summary

We are seeking a Java Spark Engineer with 7+ years of professional Java development experience and 5+ years of hands-on Apache Spark experience in production environments. The role will focus on architecting and building scalable, fault-tolerant batch and streaming data pipelines capable of processing terabyte-scale data. The ideal candidate will provide technical leadership, drive data architecture decisions, optimize distributed workloads, mentor engineers, and own production reliability while collaborating with product, analytics, and platform teams.

Key Responsibilities
• Architect and build scalable, fault-tolerant data pipelines using Apache Spark and Java.
• Lead the design and implementation of batch and streaming ETL/ELT systems handling large data volumes.
• Perform deep performance tuning, including partitioning strategies, memory management, shuffle and skew optimization, and job cost reduction.
• Establish coding standards and lead code and design reviews across the team.
• Drive technical decisions related to data architecture, storage formats, and pipeline orchestration.
• Mentor mid-level and junior engineers and serve as a technical escalation point.
• Partner with product, analytics, and platform teams to translate requirements into scalable data systems.
• Own production reliability, including on-call responsibilities, incident response, and root-cause analysis for pipeline failures.
• Evaluate and introduce new tools and frameworks where they improve system performance, reliability, or maintainability.
• Contribute to capacity planning and cost optimization for cluster infrastructure.
• Communicate technical tradeoffs, risks, and developmental challenges effectively with technical and non-technical stakeholders.
• Work effectively both independently and as part of a collaborative team in a changing environment.

Required Qualifications
• Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
• 7+ years of professional Java development experience.
• 5+ years of hands-on Apache Spark experience in production environments.
• Expert-level understanding of distributed systems, including fault tolerance, data locality, shuffle mechanics, and resource management.
• Proven experience designing systems capable of processing terabyte-scale data.
• Strong SQL skills.
• Deep familiarity with columnar and modern data storage formats, including Parquet, ORC, Avro, Delta Lake, or Iceberg.
• Experience with cluster managers such as YARN and Kubernetes and cloud-managed Spark environments.
• Proficiency with Kafka.
• Strong understanding of CI/CD, containerization, and infrastructure-as-code practices.
• Strong problem-solving and analytical skills.
• Ability to manage day-to-day technical challenges and communicate risks effectively with the technical team.

Preferred Qualifications
• Experience with Flink or other stream-processing frameworks.
• Familiarity with data governance, data lineage, and data quality frameworks.
• Experience with workflow orchestration at scale.
• Experience designing multi-tenant or multi-region data platforms.
• Prior experience leading a team or serving as a technical lead.
• Strong mentorship and coaching capabilities.
• Ability to drive ambiguous, cross-team technical initiatives.
• Strong team-oriented approach and adaptability to changing environments.

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Java Spark Engineer jobs: