Java/ Big Data so Java Spark databricks

Compunnel

$110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor’s or Master’s degree in Computer Science or related field
  • 5-8 years of software engineering expertise
  • Proven skills in Java and Python
  • Extensive experience with Apache Spark and PySpark
  • Familiarity with Data Warehousing and Data Lake architectures
  • Experience creating scalable REST APIs and microservices
  • Hands-on knowledge of AWS services and Infrastructure as Code tools

Responsibilities

  • Craft large-scale cloud-native data platforms on AWS
  • Develop distributed systems and high-performance data solutions using Spark
  • Build and maintain REST APIs and Kubernetes-based applications
  • Optimize performance, security, and cost-efficiency of data platforms
  • Implement solutions for Data Warehousing and Data Lakes
  • Utilize messaging tools like Kafka and SQS for event-driven applications
  • Lead the entire software development lifecycle from design to deployment

Benefits

  • Collaborative work environment with cross-functional teams
  • Opportunity to work with state-of-the-art technology and methodologies
  • Professional growth through exposure to diverse projects
  • Flexibility in work arrangements
  • Support for continuous learning and skill development
Full Job Description
Job Summary:

The Java / Big Data Engineer will design, develop, and deliver large-scale cloud-native data platforms primarily on AWS, leveraging Java, Python, Apache Spark, PySpark, Databricks, REST APIs, microservices, Kubernetes, and event-driven architectures. The role will work across the technology stack to build highly scalable, resilient, secure, and cost-efficient systems while owning the end-to-end software development lifecycle and collaborating with engineering, product, architecture, and business stakeholders.

Key Responsibilities:
• Deliver large-scale cloud-native data platforms primarily on AWS using REST APIs, microservices, and event-driven applications.
• Work hands-on across Java, Python, Spark, PySpark, TypeScript, JavaScript, Angular, AWS services, event-driven architectures, and SQL/NoSQL databases.
• Design and develop scalable distributed systems and high-performance data processing solutions using Apache Spark and PySpark.
• Develop and maintain REST APIs, microservices, and Kubernetes-based containerized applications.
• Drive performance optimization, scalability, reliability, security, governance, and cost efficiency across data platforms.
• Design and implement solutions involving Data Warehousing, Data Lakes, and Delta Lake architectures.
• Develop cloud-native data platforms using AWS services such as S3, Lambda, API Gateway, and EventBridge.
• Implement messaging and event-driven solutions using Kafka, SNS, and SQS.
• Work with relational and NoSQL databases including PostgreSQL, SQL Server, Aurora, DynamoDB, MongoDB, and Redis.
• Implement Infrastructure as Code using Terraform and Ansible.
• Support CI/CD and DevOps practices using Jenkins, GitHub/GitLab, Bitbucket, GoCD, and automated deployment pipelines.
• Implement unit, integration, and regression testing strategies while following engineering best practices.
• Own the end-to-end software development lifecycle, including requirements gathering, solution design, development, deployment, observability, and documentation.
• Collaborate with global engineering, product management, architecture, and business stakeholders to align technical solutions with business objectives.
• Diagnose, troubleshoot, and resolve complex technical problems using strong critical thinking and analytical skills.
• Apply awareness of Generative AI technologies, including LLMs, RAG architectures, and Agentic AI systems, where applicable.

Required Qualifications:
• B.E., B.Tech, M.Tech, MCA, or equivalent degree in Computer Science, Information Technology, or a related field.
• 5-8 years of strong software engineering experience building scalable applications and distributed systems.
• Strong hands-on expertise in Java and Python.
• Strong experience with Apache Spark and PySpark for distributed data processing.
• Experience with Data Warehousing, Data Lakes, Delta Lake architecture, and modern Big Data ecosystem designs.
• Proven experience designing and building scalable REST APIs, microservices, Kubernetes-based applications, and distributed systems.
• Strong understanding of object-oriented design patterns and functional programming.
• Experience with AWS services including S3, Lambda, API Gateway, and EventBridge.
• Strong experience with Kafka, SNS, and SQS and event-driven architectures.
• Experience with relational and NoSQL databases such as PostgreSQL, SQL Server, Aurora, DynamoDB, MongoDB, and Redis.
• Hands-on experience with Terraform and Ansible for Infrastructure as Code.
• Understanding of CI/CD and DevOps practices and automated deployment pipelines.
• Experience implementing unit, integration, and regression testing strategies.
• Strong critical thinking, analytical, troubleshooting, and problem-solving skills.
• Awareness of Generative AI technologies, including LLMs, RAG architectures, and Agentic AI systems.

Preferred Qualifications:
• Working knowledge of PySpark with Databricks.
• Experience working with Azure and/or Google Cloud Platform (GCP).
• Experience building data platforms in privacy-safe, Customer Data Platform, or Marketing Technology environments.
• Experience with Angular, TypeScript, and JavaScript in UX-driven applications.

Similar Jobs

More Jobs at Compunnel

More Information Technology Jobs

Find similar Java/ Big Data so Java Spark databricks jobs: