Job Summary
We are seeking a Big Data Engineer with deep expertise in AWS EKS and Kubernetes to support large-scale Apache Spark workloads. The role requires a subject matter expert who can advise on best practices and architecture for running and optimizing Spark workloads on Kubernetes. The ideal candidate will have experience working in AWS or large enterprise cloud environments and with Spark at scale on EKS, including hands-on experience with EMR on EKS and petabyte-scale data processing environments.
Key Responsibilities
• Design, support, and optimize large-scale Apache Spark workloads running on Kubernetes.
• Provide subject matter expertise and architectural guidance for AWS EKS and Spark-on-Kubernetes environments.
• Work with EMR on EKS to deploy, manage, and optimize Spark workloads.
• Apply best practices across Kubernetes architecture, workload management, resource allocation, and autoscaling.
• Manage and troubleshoot Kubernetes pods, deployments, namespaces, and services.
• Work with Helm for Kubernetes application deployment and management.
• Address Kubernetes networking and security requirements.
• Troubleshoot cluster and workload performance issues and implement appropriate improvements.
• Develop and support Big Data solutions using Spark, Hadoop, Hive, and Trino.
• Work with AWS services including EKS, EMR, S3, Glue, Athena, and Lambda.
• Develop data engineering solutions using Python and SQL, with Scala preferred.
• Support petabyte-scale data processing environments.
• Provide guidance on best practices and architecture for large-scale cloud-based data processing environments.
Required Qualifications
• Proven expertise with AWS EKS (Elastic Kubernetes Service).
• Hands-on experience running and optimizing Apache Spark workloads on Kubernetes.
• Hands-on experience with EMR on EKS.
• Strong understanding of Kubernetes architecture, including pods, deployments, namespaces, and services.
• Strong understanding of resource allocation and autoscaling.
• Experience with Helm.
• Strong understanding of Kubernetes networking and security.
• Experience troubleshooting Kubernetes cluster and workload performance issues.
• Strong experience with Big Data technologies including Spark, Hadoop, Hive, and Trino.
• Experience with AWS services including EKS, EMR, S3, Glue, Athena, and Lambda.
• Strong coding experience with Python and SQL.
• Experience working with petabyte-scale data processing environments.
• Proven large-scale Big Data engineering experience.
• Relevant AWS and/or Kubernetes certifications.
Preferred Qualifications
• Experience with Scala.
• Experience with Terraform or CloudFormation.
• Experience with CI/CD tools such as Jenkins, GitLab, GitHub Actions, or ArgoCD.
• Experience with Docker.
• Experience with Prometheus and Grafana.
• Experience with Service Mesh technologies.
• Financial Services industry experience.
• Experience with GenAI tools such as Copilot, ChatGPT, Claude, or Amazon Q.