EPAM Systems

Senior Data Software Engineer, Spark, Java, Scala

EPAM Systems$120K — $145K *
US-AnywhereRemote in Georgia, US
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of hands-on big-data experience, particularly with Hadoop and Spark
  • Excellent SQL skills (Spark SQL, HiveQL) for Big Data contexts
  • Strong knowledge of Spark, including optimization techniques for execution plans
  • Proficient in Scala or Java programming
  • Understanding of ETL principles and batch processing in Data Warehouses
  • Familiar with data completeness signals and orchestration methods
  • Strong English communication skills, both verbal and written

Responsibilities

  • Write data pipelines and jobs for new outputs related to feature development
  • Integrate existing pipelines with new organizational tools and services
  • Troubleshoot and fix bugs in pipelines and data stemming from logic errors
  • Perform ad-hoc data exploration and validation to support tech design decisions
  • Monitor and address production pipeline issues
  • Develop and implement data quality checks for system monitoring
  • Scope and plan new development efforts, providing timelines and assessments

Benefits

  • Opportunities for professional growth and development
  • Collaborative work environment with cross-functional teams
  • Focus on high-quality, maintainable code alongside data integrity
  • Engagement in innovative projects shaping data as a core product
  • Access to cutting-edge tools and technologies used in big data processing
Full Job Description
We are seeking a Senior Software Data Engineer to join our team in a software engineering capacity. This is not a data-science or analytics position centered on ad-hoc data exploration; instead, the role focuses on building software, data processing jobs, and data pipelines consumed by internal and external partners. Data is our main product and first-class citizen, and we value correct, high-quality data as much as clean and maintainable code. Responsibilities Write new data pipelines and jobs to produce new outputs (datasets) in scope of new features development Adopt existing data pipelines to integrate with new org-wide platforms, tools, services, and languages Fix bugs in code and correct data caused by incorrect logic or implementation Perform ad-hoc data exploration, validation, and investigation to help select the right tech design and support Product Management team decisions Monitor and troubleshoot production issues with pipelines owned by the team Develop and adopt data quality checks to monitor data issues in the systems Scope and plan new development, including assessing level of effort and providing timelines Maintain tickets hygiene in Radar (ticketing system) Evolve jobs, apps, and systems to a better state across all aspects: code quality, complexity, maintainability, and documentation Communicate with other data engineers in the team, peer teams (QA, UAT, Platform, etc), project managers, and engineering managers on status, blockers, estimates, and timelines Requirements 3+ years of hands-on experience in the big-data field, including Hadoop (HDFS, YARN or Mesos) and Spark Excellent knowledge and hands-on experience of SQL in context of Big Data: Spark SQL, HiveQL Excellent knowledge of Spark, including ability to understand and optimize Spark execution plans via Spark UI, with upcoming migration to Spark 3 Excellent knowledge of Scala or Java Understanding of batch processing and ETL principles in Data Warehouses Familiarity with data completeness signals and orchestration Knowledge of approaches for historical reprocessing and data correction Skills in handling bad data and late data in inputs and outputs Understanding of schema migrations and datasets evolution Strong speaking English, with ability to rely on information heard verbally in meetings and to explain own ideas clearly to native speakers Capability to learn fast new set of tools and technology used internally at the company: platform services, telemetry providers, Spark-as-a-Service, build system, and more Nice to have Understanding of functional programming ideas and principles Experience in building and using web services Familiarity with any of Teradata, Vertica, Oracle, Tableau Skills in Spark Streaming and Kafka Knowledge of Apache Iceberg, Trino (Presto), Druid, Cassandra, or Blob storage like AWS Experience with Splunk Experience with Snowflake

About EPAM Systems

EPAM Systems, Inc. is a leading global provider of digital platform engineering and development services. The company has a strong presence in North America, Europe, and Asia, and serves clients in a variety of industries, including financial services, healthcare, and retail. EPAM's services include software engineering, product development, and digital platform engineering, and the company has a reputation for delivering high-quality solutions that help its clients achieve their business goals. EPAM has been recognized as a leader in the digital services industry by a number of independent research firms, and the company has won numerous awards for its work.
Learn more about EPAM Systems
Size
58,824 employees
Market Cap
$18.2 billion
Industry
Net Income
$327.1 million
Founded
1993
5 Year Trend
+26.5%
Revenue
$2.6 billion
NASDAQ

Similar Jobs

More Jobs at EPAM Systems

More Enterprise Technology Jobs

Find similar Senior Data Software Engineer, Spark, Java, Scala jobs: