Kafka/Spark Developer

System One Holdings, LLC

• $110K — $130K *
Information Technology
5 - 7 years of experience
Job Overview by Ladders

Qualifications

  • 5+ years of experience in Big Data development and distributed data processing environments.
  • Strong hands-on experience with Apache Kafka and its configurations.
  • Extensive capability in real-time data processing with Apache Spark Streaming.
  • Proficiency in programming languages such as Java, Scala, or Python (PySpark).
  • Skills in writing and optimizing SQL queries using Impala or similar engines.
  • Experience with Hadoop ecosystem components like HDFS and Hive.
  • Strong analytical and troubleshooting skills in streaming environments.

Responsibilities

  • Develop scalable big data solutions with Hadoop, Spark, Kafka, and Impala.
  • Design and optimize data pipelines for batch and real-time processing.
  • Create Spark applications for data transformation and analysis.
  • Maintain Kafka producers and consumers for real-time data flows.
  • Implement logical and physical data models for reporting and analytics.
  • Monitor and tune Kafka and Spark jobs for performance improvements.
  • Collaborate with cross-functional teams to build modern data platforms.

Benefits

  • Hybrid work model, offering both onsite and remote flexibility.
  • Opportunity to work with cutting-edge big data technologies.
  • Collaborative team environment with Agile methodologies.
  • Professional development and growth opportunities.
Full Job Description
Job Title: Kafka/Spark Developer
Location: Pittsburgh, Pennsylvania
Type: Permanent Full Time
Work Model: Hybrid - onsite and remote

Responsibilities

  • Develop and maintain scalable big data solutions using Hadoop, Spark, Kafka, and Impala to support enterprise data processing and analytics initiatives.
  • Design, build, and optimize batch and real-time data pipelines for ingesting, processing, transforming, and delivering large volumes of structured and unstructured data.
  • Develop Spark applications using PySpark, Scala, or Java for data transformation, aggregation, cleansing, and analytical processing.
  • Build and maintain Kafka producers, consumers, topics, and streaming workflows to enable reliable real-time data ingestion and event-driven architectures.
  • Design and implement logical and physical data models to support data warehousing, reporting, analytics, and business intelligence requirements.
  • Monitor, troubleshoot, and tune Kafka and Spark streaming jobs to improve performance, scalability, and operational reliability.
  • Optimize Hadoop ecosystem components, Spark jobs, Kafka configurations, and Impala queries to improve system performance and resource utilization.
  • Collaborate with architects, data engineers, DevOps teams, and business stakeholders to design and implement modern streaming and event-driven data platforms.
  • Analyze user requirements, and define technical project scope and assumptions for assigned tasks.
  • Create technical designs for new systems, and/or modifications to existing systems.
  • Translate detailed requirements into functional system designs.
  • Prioritize work, meet deadlines, and establish and maintain effective working relationships with clients, project team members, supervisors, and employees from other departments.
  • Partner with business leaders, enterprise architects, and product owners to identify new graph-based use cases, evaluate emerging technologies, and align Neo4j initiatives with digital transformation goals.


Requirements

  • At least 5+ years of experience in Big Data development, data engineering, or distributed data processing environments.
  • Strong hands-on experience with Apache Kafka, topic configuration, producer/consumer development, Kafka Connect, and Schema Registry.
  • Extensive experience developing real-time data processing applications using Apache Spark Streaming and/or Spark Structured Streaming.
  • Proficiency in Java, Scala, or Python (PySpark) with strong object-oriented programming and software development skills.
  • Proficiency in writing and optimizing complex SQL queries using Impala, Hive, or similar distributed query engines.
  • Hands-on experience with Hadoop ecosystem components including HDFS, Hive.
  • Experience integrating Kafka and Spark with relational databases, NoSQL databases, cloud storage platforms, and enterprise applications.
  • Strong analytical, troubleshooting, and performance tuning skills in distributed streaming environments.
  • Excellent communication, collaboration, and stakeholder management skills, with the ability to work effectively in Agile/Scrum teams.
  • Experience working in Agile development environments with strong collaboration, technical leadership, problem-solving, and stakeholder communication skills.


#M-
#LI-
Ref: #404-IT Pittsburgh

Similar Jobs

More Jobs at System One Holdings, LLC

More Information Technology Jobs

Find similar Kafka/Spark Developer jobs: