Hitachi America

Sr. Data Engineer - Flink

Hitachi America$120K — $145K *
Information Technology
8 - 10 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, or related field or equivalent experience.
  • 10+ years of experience in software and data engineering, with 7+ years focused on Apache Flink.
  • Deep familiarity with AWS Managed Service for Apache Flink and production scaling.
  • Proficient in Flink DataStream API and SQL, particularly in stateful processing.
  • Experience diagnosing performance issues in streaming platforms including backpressure and memory problems.
  • Strong skills in Java, Python/PyFlink, and SQL for production streaming applications.
  • Understanding of Kafka, including topic designs and sharded distributed systems.

Responsibilities

  • Lead and mentor data engineering teams while engaging in hands-on solution development.
  • Act as the technical authority on Kafka, Apache Flink, and AWS managed services for streaming applications.
  • Oversee architecture and production readiness reviews for data solutions.
  • Define and implement engineering standards and best practices for streaming data.
  • Troubleshoot complex production issues and lead root-cause analysis efforts.
  • Document procedures and operational guidelines for ongoing support and deployments.
  • Develop production-scale streaming architectures using stateful processing and fault tolerance strategies.

Benefits

  • Opportunities for professional development and training programs.
  • Flexible work environment and work-life balance initiatives.
  • Collaborative team culture focused on innovation and excellence.
  • Comprehensive health, dental, and vision insurance packages.
Full Job Description

Function

Cloud & Data Engineering

Job description

Job Summary

We are seeking a highly skilled Hands-On Data Engineering Lead with deep expertise in Apache Flink, AWS Managed Service for Apache Flink, Kafka, and AWS. This role requires a technical leader who actively architects, develops, debugs, optimizes, and supports production-grade real-time streaming platforms. The ideal candidate combines hands-on engineering depth with experience leading teams that deliver scalable, resilient, high-throughput, and low-latency data solutions in 24x7 environments.

Key Responsibilities

Hands-On Technical Leadership

  • Lead and mentor data engineers and architects while remaining directly involved in solution design, implementation, and troubleshooting.
  • Serve as the technical authority for Apache Flink, AWS Managed Service for Apache Flink, Kafka-based streaming, and associated AWS services.
  • Lead architecture, design, code, configuration, deployment, and production-readiness reviews.
  • Establish engineering standards, coding practices, test discipline, and production-support procedures.
  • Coordinate technical decisions across application, data, cloud-platform, performance, and operations teams.

Real-Time Streaming Architecture & Engineering

  • Architect, design, and develop production-scale streaming solutions using Apache Flink, AWS Managed Service for Apache Flink, Java or PyFlink, Flink SQL, and Kafka.
  • Apply stateful and event-time processing patterns, including keyed and broadcast state, state TTL, timers, windows, joins, watermarks, and delivery-semantics controls.
  • Design resilient checkpointing, savepoint, restart, recovery, and state-migration strategies.
  • Design Kafka topics, partitions, shard keys, routing, consumer groups, offsets, transactions, schemas, and source/sink integrations.
  • Design and operate sharded streaming architectures, including workload decomposition, shard-key selection, state distribution, rebalancing, cross-shard processing, parallelism, failure isolation, and recovery.
  • Build fault-tolerant streaming applications that meet defined throughput, latency, availability, and recovery objectives.

Performance Engineering & Optimization

  • Define measurable entry, exit, and acceptance criteria for load, peak, burst, replay, recovery, and long-running soak tests.
  • Verify that test inputs, replay behavior, duration, measurement windows, metric units, and outputs are representative and comparable.
  • Analyze throughput, latency, backpressure, state growth, checkpoints, recovery, resource utilization, data skew, and operating headroom.
  • Identify operator and stage-level bottlenecks and implement validated code, configuration, partitioning, or scaling improvements.
  • Document test conditions, findings, qualifications, risks, and recommendations using reproducible evidence.

Production Engineering & Troubleshooting

  • Act as a senior escalation point for complex production issues involving Apache Flink, AWS Managed Service for Apache Flink, and Kafka.
  • Diagnose checkpoint failures, savepoint recovery issues, backpressure, state growth, idle partitions, data skew, memory pressure, garbage collection, restarts, and throughput or latency degradation.
  • Analyze JobManager and TaskManager events, runtime configuration, logs, metrics, deployment history, and service behavior.
  • Lead evidence-based root-cause analysis and define corrective and preventive actions for material incidents.
  • Use controlled experiments to confirm or reject technical hypotheses and validate remediation effectiveness.
  • Engage AWS Support and service specialists when deeper platform analysis or service-limit clarification is required.

Observability & Operational Readiness

  • Define and implement monitoring for throughput, lag, backpressure, state size, checkpoints, restarts, failures, service events, and recovery.
  • Standardize metric definitions, units, aggregation windows, data sources, thresholds, alerting, and escalation paths.
  • Develop and review deployment, incident-triage, rollback, snapshot/savepoint recovery, scaling, and change-control procedures.
  • Prepare technical documentation, operational runbooks, and knowledge-transfer materials for engineering and support teams.

Software Engineering Excellence

  • Write, debug, test, optimize, and deploy production-grade Java, Python/PyFlink, and SQL code.
  • Perform code reviews and promote automated testing, code quality, infrastructure-as-code, and DevOps practices.
  • Build and improve CI/CD pipelines and deployment automation for streaming applications.

Required Qualifications

  • Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.
  • 10+ years of software engineering, data engineering, or distributed-systems experience, including 7+ years of recent hands-on Apache Flink engineering.
  • Deep production experience with AWS Managed Service for Apache Flink, including deployments, runtime behavior, scaling, service limits, snapshots, monitoring, troubleshooting, and recovery.
  • Strong command of the Flink DataStream API and Flink SQL, including stateful processing, event time, watermarks, windows, joins, timers, checkpoints, savepoints, and exactly-once concepts.
  • Demonstrated experience diagnosing backpressure, checkpoint behavior, state growth, skew, idle partitions, memory and GC issues, restarts, Kafka lag, and performance degradation.
  • Strong programming skills in Java, Python/PyFlink, and SQL, with experience developing and supporting production-grade streaming applications.
  • Hands-on experience with Kafka or Confluent Kafka and sharded distributed systems, including topic and shard-key design, partitioning, routing, consumer groups, rebalancing, state migration, cross-shard processing, failure isolation, and recovery.
  • Working knowledge of AWS services and controls supporting streaming platforms, including CloudWatch, S3, IAM, networking, service quotas, CI/CD, and infrastructure as code.
  • Experience building and operating highly available, high-throughput, low-latency platforms in 24x7 production environments.
  • Experience with monitoring, alerting, observability, incident response, root-cause analysis, and recovery planning.
  • Ability to communicate technical findings clearly and distinguish observed facts, estimates, hypotheses, proposals, and approved decisions.
  • Demonstrated ability to lead technical teams while remaining hands-on in engineering and troubleshooting activities.

Preferred Qualifications

  • Experience supporting global, mission-critical streaming platforms and mentoring engineering teams.
  • AWS certification or demonstrated equivalent AWS platform expertise.
  • Experience processing very large event volumes using multi-shard architectures and managing multi-terabyte state, workload skew, state migration, or disaster-recovery design.

About Hitachi America

Hitachi America is a subsidiary of Hitachi, Ltd., a Japanese multinational conglomerate. They provide a wide range of products and services, including information technology, power systems, and social infrastructure. They work with clients in a variety of industries, including healthcare, transportation, and finance. They are committed to sustainability and social responsibility, and have implemented various initiatives to reduce their environmental impact.
Learn more about Hitachi America
Size
368,247 employees
Industry
Founded
1959
NASDAQ

Similar Jobs

More Jobs at Hitachi America

More Information Technology Jobs

Find similar Sr. Data Engineer - Flink jobs: