EPAM Systems

Senior Data Engineer/ Databricks, PySpark, Python

EPAM Systems$125K — $150K *
US-AnywhereRemote in Georgia, US
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of hands-on experience in data engineering
  • Expertise in Azure Databricks
  • Proficiency in PySpark and distributed data processing
  • Advanced skills in Python development
  • Understanding of performance optimization for large-scale Spark workloads

Responsibilities

  • Monitor the health, stability, and performance of the production data platform
  • Investigate and resolve performance issues and production incidents
  • Analyze Databricks workloads and identify performance bottlenecks
  • Optimize PySpark workloads and Databricks jobs for performance, reliability, and cost efficiency
  • Support and improve approximately 200 existing Databricks jobs and associated data pipelines
  • Design and implement scaling strategies with a strong focus on predictable infrastructure and operational costs
  • Track and optimize Databricks and Azure resource consumption

Benefits

  • Work on a high-impact project with a major European automotive manufacturer
  • Opportunity to operate in a multi-regional deployment environment
  • Engage in the evolution phase of a cutting-edge data platform
  • Potential for professional growth in cloud and data engineering practices
  • Collaborate with an experienced team of data professionals
Full Job Description
We are looking for an experienced Senior Data Engineer / Databricks Engineer to join a project focused on collecting, processing, and analyzing vehicle telemetry data for a major European automotive manufacturer. The platform operates at a significant scale, processing tens of terabytes of telemetry data in Azure Databricks and running approximately 200 Databricks jobs. The solution is currently in the production maintenance and evolution phase and is deployed across multiple regions, including Europe, North America, and China, with an additional deployment in China planned in the near future. A key focus of the role will be ensuring that the system can scale reliably while keeping infrastructure and operational costs predictable. Responsibilities Monitor the health, stability, and performance of the production data platform Investigate and resolve performance issues and production incidents Analyze Databricks workloads and identify performance bottlenecks Optimize PySpark workloads and Databricks jobs for performance, reliability, and cost efficiency Support and improve approximately 200 existing Databricks jobs and associated data pipelines Prepare the platform for increasing data volumes and workloads as additional vehicle platforms are onboarded Design and implement scaling strategies with a strong focus on predictable infrastructure and operational costs Track and optimize Databricks and Azure resource consumption Improve platform observability, monitoring, alerting, and operational processes Ensure consistency across multiple regional deployments in Europe, North America, and China Identify opportunities for technical improvements, automation, and reduction of operational overhead Requirements 3+ years of hands-on experience in data engineering Expertise in Azure Databricks Proficiency in PySpark and distributed data processing Advanced skills in Python development Understanding of performance optimization for large-scale Spark workloads Background in operating and troubleshooting production data platforms Knowledge of monitoring, observability, and capacity planning combined with cloud cost optimization Ability to analyze existing systems, identify bottlenecks, and propose pragmatic improvements English proficiency at an Upper-Intermediate level (B2) or higher Nice to have Familiarity with Microsoft Azure services and cloud infrastructure Experience building and maintaining CI/CD pipelines using GitHub Actions Exposure to large-scale telemetry, IoT, or automotive data Background in managing Databricks environments with a large number of scheduled jobs and pipelines Showcase of multi-region or geographically distributed cloud deployments

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Information Technology Jobs

Find similar Senior Data Engineer/ Databricks, PySpark, Python jobs: