Function
Cloud & Data Engineering
Job descriptionJob Summary
We are seeking a highly skilled Hands-On Data Engineering Lead with deep expertise in Apache Flink, AWS Managed Service for Apache Flink, Kafka, and AWS. This role requires a technical leader who actively architects, develops, debugs, optimizes, and supports production-grade real-time streaming platforms. The ideal candidate combines hands-on engineering depth with experience leading teams that deliver scalable, resilient, high-throughput, and low-latency data solutions in 24x7 environments.
Key Responsibilities
Hands-On Technical Leadership
- Lead and mentor data engineers and architects while remaining directly involved in solution design, implementation, and troubleshooting.
- Serve as the technical authority for Apache Flink, AWS Managed Service for Apache Flink, Kafka-based streaming, and associated AWS services.
- Lead architecture, design, code, configuration, deployment, and production-readiness reviews.
- Establish engineering standards, coding practices, test discipline, and production-support procedures.
- Coordinate technical decisions across application, data, cloud-platform, performance, and operations teams.
Real-Time Streaming Architecture & Engineering
- Architect, design, and develop production-scale streaming solutions using Apache Flink, AWS Managed Service for Apache Flink, Java or PyFlink, Flink SQL, and Kafka.
- Apply stateful and event-time processing patterns, including keyed and broadcast state, state TTL, timers, windows, joins, watermarks, and delivery-semantics controls.
- Design resilient checkpointing, savepoint, restart, recovery, and state-migration strategies.
- Design Kafka topics, partitions, shard keys, routing, consumer groups, offsets, transactions, schemas, and source/sink integrations.
- Design and operate sharded streaming architectures, including workload decomposition, shard-key selection, state distribution, rebalancing, cross-shard processing, parallelism, failure isolation, and recovery.
- Build fault-tolerant streaming applications that meet defined throughput, latency, availability, and recovery objectives.
Performance Engineering & Optimization
- Define measurable entry, exit, and acceptance criteria for load, peak, burst, replay, recovery, and long-running soak tests.
- Verify that test inputs, replay behavior, duration, measurement windows, metric units, and outputs are representative and comparable.
- Analyze throughput, latency, backpressure, state growth, checkpoints, recovery, resource utilization, data skew, and operating headroom.
- Identify operator and stage-level bottlenecks and implement validated code, configuration, partitioning, or scaling improvements.
- Document test conditions, findings, qualifications, risks, and recommendations using reproducible evidence.
Production Engineering & Troubleshooting
- Act as a senior escalation point for complex production issues involving Apache Flink, AWS Managed Service for Apache Flink, and Kafka.
- Diagnose checkpoint failures, savepoint recovery issues, backpressure, state growth, idle partitions, data skew, memory pressure, garbage collection, restarts, and throughput or latency degradation.
- Analyze JobManager and TaskManager events, runtime configuration, logs, metrics, deployment history, and service behavior.
- Lead evidence-based root-cause analysis and define corrective and preventive actions for material incidents.
- Use controlled experiments to confirm or reject technical hypotheses and validate remediation effectiveness.
- Engage AWS Support and service specialists when deeper platform analysis or service-limit clarification is required.
Observability & Operational Readiness
- Define and implement monitoring for throughput, lag, backpressure, state size, checkpoints, restarts, failures, service events, and recovery.
- Standardize metric definitions, units, aggregation windows, data sources, thresholds, alerting, and escalation paths.
- Develop and review deployment, incident-triage, rollback, snapshot/savepoint recovery, scaling, and change-control procedures.
- Prepare technical documentation, operational runbooks, and knowledge-transfer materials for engineering and support teams.
Software Engineering Excellence
- Write, debug, test, optimize, and deploy production-grade Java, Python/PyFlink, and SQL code.
- Perform code reviews and promote automated testing, code quality, infrastructure-as-code, and DevOps practices.
- Build and improve CI/CD pipelines and deployment automation for streaming applications.
Required Qualifications
- Bachelor's or Master's degree in Computer Science, Engineering, Information Systems, or a related field, or equivalent professional experience.
- 10+ years of software engineering, data engineering, or distributed-systems experience, including 7+ years of recent hands-on Apache Flink engineering.
- Deep production experience with AWS Managed Service for Apache Flink, including deployments, runtime behavior, scaling, service limits, snapshots, monitoring, troubleshooting, and recovery.
- Strong command of the Flink DataStream API and Flink SQL, including stateful processing, event time, watermarks, windows, joins, timers, checkpoints, savepoints, and exactly-once concepts.
- Demonstrated experience diagnosing backpressure, checkpoint behavior, state growth, skew, idle partitions, memory and GC issues, restarts, Kafka lag, and performance degradation.
- Strong programming skills in Java, Python/PyFlink, and SQL, with experience developing and supporting production-grade streaming applications.
- Hands-on experience with Kafka or Confluent Kafka and sharded distributed systems, including topic and shard-key design, partitioning, routing, consumer groups, rebalancing, state migration, cross-shard processing, failure isolation, and recovery.
- Working knowledge of AWS services and controls supporting streaming platforms, including CloudWatch, S3, IAM, networking, service quotas, CI/CD, and infrastructure as code.
- Experience building and operating highly available, high-throughput, low-latency platforms in 24x7 production environments.
- Experience with monitoring, alerting, observability, incident response, root-cause analysis, and recovery planning.
- Ability to communicate technical findings clearly and distinguish observed facts, estimates, hypotheses, proposals, and approved decisions.
- Demonstrated ability to lead technical teams while remaining hands-on in engineering and troubleshooting activities.
Preferred Qualifications
- Experience supporting global, mission-critical streaming platforms and mentoring engineering teams.
- AWS certification or demonstrated equivalent AWS platform expertise.
- Experience processing very large event volumes using multi-shard architectures and managing multi-terabyte state, workload skew, state migration, or disaster-recovery design.