Data Engineer (Temporal & Apache Kafka required)

Infinitive Inc

$90K — $154K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 4+ years in data engineering or backend systems
  • Hands-on with Temporal or Cadence for workflow orchestration
  • Expertise in Apache Kafka event streaming
  • Proficient with schema design frameworks like Avro and Protobuf
  • Strong Python programming skills; Go or Java knowledge is a plus
  • Experience with Apache Spark for large-scale data processing
  • Advanced SQL skills for data modeling and performance tuning

Responsibilities

  • Architect and manage distributed data streams with Apache Kafka
  • Establish schema design standards and automated validation
  • Implement durable workflows using Temporal for data pipelines
  • Build end-to-end pipelines using Python and SQL
  • Design analytical data models in cloud-based data warehouses
  • Ensure data quality and reliability through automated testing
  • Collaborate with cross-functional teams to define data contracts

Benefits

  • Comprehensive health insurance packages
  • Flexible remote work options
  • Professional development and training opportunities
  • Collaborative and innovative work environment
  • Retirement savings plan with employer matching
Full Job Description
About the Role

We are seeking an experienced Data Engineer to help design, build, and scale our next-generation event-driven data platforms. In this role, you will be instrumental in bridging high-throughput distributed streaming with complex, fault-tolerant workflow orchestration and strict data governance.

You will work extensively with Apache Kafka for real-time event streaming and Temporal (the open-source, durable execution engine originating from Uber/Cadence) to build resilient, distributed stateful workflows and data pipelines. A core focus of this position is establishing robust data schema design and automated validation to ensure strong data contracts across distributed systems. Alongside these technologies, you will design robust batch and streaming ETL/ELT pipelines leveraging Python, Apache Spark, and modern cloud data warehouses/lakehouses.Key Responsibilities
  • Stream Processing & Messaging: Architect, deploy, and maintain high-volume distributed data streams using Apache Kafka (producers, consumers, Kafka Connect, Schema Registry).
  • Data Schema Design & Validation: Establish and enforce schema design standards, versioning strategies, and automated schema validation (e.g., Avro, Protobuf, JSON Schema) to maintain strict data contracts across microservices, streaming consumers, and lakehouse storage.
  • Resilient Workflow Orchestration: Design and implement durable execution workflows using Temporal to coordinate long-running distributed pipelines, compensate transactions (Saga pattern), and manage cross-system ETL tasks.
  • Pipeline Development: Build end-to-end batch and near-real-time pipelines using Python, SQL, and Apache Spark / PySpark.
  • Data Modeling & Warehousing: Design and optimize analytical data models (dimensional/star schema) in modern cloud data warehouses/lakehouses (e.g., Snowflake, BigQuery, Databricks, Redshift).
  • Reliability & Data Quality: Implement automated testing, continuous schema validation, data drift detection, and observability across streaming and batch workflows.
  • Cross-Functional Collaboration: Partner with software engineers, machine learning engineers, and analysts to define standard schema definitions, data contracts, and production-grade CI/CD release patterns.

Required Experience
  • 4+ years of professional experience in data engineering, backend distributed systems, or software engineering.
  • Hands-on experience with Temporal (or Cadence): Proven understanding of durable workflows, activities, retries, signals, queries, and long-running distributed task orchestration.
  • Deep expertise with Apache Kafka: Practical experience with message partitioning, consumer groups, offset management, and topic design.
  • Strong background in Data Schema Design & Validation:
    • Demonstrated proficiency with schema definition frameworks (Apache Avro, Protocol Buffers/gRPC, or JSON Schema).
    • Practical experience managing schema evolution, compatibility modes (backward/forward/full), and schema registries (e.g., Confluent Schema Registry, AWS Glue Schema Registry).
    • Experience enforcing data validation rules, contract testing, and data quality checks (e.g., Great Expectations, Pandera, Pydantic, dbt tests).
  • Strong programming proficiency in Python (Go or Java is a plus) with clean code, design patterns, and unit/integration testing standards.
  • Distributed computing experience: Hands-on development with Apache Spark (PySpark/Spark SQL) processing large-scale datasets.
  • Advanced SQL & Data Modeling: Strong experience with relational databases, dimensional data modeling, and query performance tuning.
Infinitive is required by law in some jurisdictions to include a reasonable estimate of the compensation range for this role. The determination of this range includes various factors not limited to skill set, level, experience, relevant training, and licensure and certifications. Compensation decisions are dependent on the facts and circumstances of each case. A reasonable estimate of the current range for this role in the U.S. is $90,000 - $154,00.00.

Similar Jobs

More Jobs at Infinitive Inc

More Information Technology Jobs

Find similar Data Engineer (Temporal & Apache Kafka required) jobs: