Data Engineer, AI & Distributed Systems

Zignal Labs

$120K — $140K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of experience building and operating data pipelines in production.
  • Strong programming skills in Scala, Java, or Kotlin; Scala preferred.
  • Working proficiency in Python for data and scripting.
  • Hands-on experience with distributed processing frameworks like Apache Spark.
  • Familiarity with streaming platforms, particularly Kafka, including concepts like consumer groups and offsets.
  • Practical experience with AWS and Docker in a Kubernetes environment.
  • Knowledge of a workflow orchestrator like Airflow, Prefect, or Dagster.
  • Solid SQL skills and experience with NoSQL or caching solutions.

Responsibilities

  • Build and maintain data pipelines for high-volume unstructured data ingestion.
  • Support AI systems by developing data pathways for NLP, LLM, and retrieval services.
  • Integrate and optimize search and vector stores for semantic search and real-time retrieval.
  • Implement and improve microservices and APIs for enterprise analytics delivery.
  • Participate in code reviews, CI/CD, and improve production reliability through debugging.
  • Collaborate with cross-functional teams to transition prototypes into production.

Benefits

  • Fully remote work environment with flexible scheduling.
  • Opportunity for professional growth with a focus on skill development and learning.
  • Access to cutting-edge technologies in data processing and streaming.
  • Collaborative culture that emphasizes team input and idea-sharing.
  • Engagement in innovative projects that impact real-time intelligence applications.
Full Job Description
About the Role

We ingest, enrich, and structure massive volumes of unstructured data - from social platforms and news outlets to broadcast media - and turn it into real-time intelligence for our customers.

As a Data Engineer on this team, you'll build and operate the pipelines that make that possible. You'll work on systems that process billions of events a day, and on the data pathways that feed our search, NLP, and AI services. You'll own meaningful pieces of the pipeline end to end, and you'll do it alongside engineers who have been running these systems at scale for years.

This is a hands-on build-and-operate role. You don't need to have designed a distributed system from scratch before - you need to be someone who writes solid code, reasons carefully about data correctness and failure modes, and wants to go deep on streaming and AI infrastructure.
What You'll Do
  • Build and maintain pipelines. Develop and operate batch and streaming pipelines that ingest and enrich high-volume unstructured data. Own components end to end, from implementation through production monitoring.
  • Support our AI systems. Build and extend the data pathways that feed downstream NLP, LLM, and retrieval services - including data preparation, embedding generation, and indexing workflows.
  • Work with search and storage layers. Integrate with and tune our search and vector stores to support semantic search, clustering, and real-time retrieval.
  • Build services and APIs. Implement and improve the microservices and APIs that deliver analytics to enterprise customers.
  • Operate what you build. Write clean, tested, maintainable code. Participate in code review, CI/CD, and infrastructure-as-code practices. Debug production issues and improve reliability over time.
  • Collaborate across teams. Work with Data Science, ML, Product, and Security to take ideas from prototype into production.
What You'll Need

These are the things we genuinely need on day one.
  • 3+ years building and operating data pipelines in production.
  • Strong programming skills in a JVM language - Scala, Java, or Kotlin. Our core pipeline code is Scala. If you're strong in Java or Kotlin and want to learn Scala, we'll support that; we care more about your fundamentals than your current syntax.
  • Working proficiency in Python for data and scripting work.
  • Hands-on experience with a distributed processing framework, most likely Apache Spark.
  • Hands-on experience with a streaming platform, most likely Kafka - including a real understanding of consumer groups, offsets, partitioning, and what happens when things fall behind.
  • Practical AWS experience and comfort with Docker. You should be able to work in a Kubernetes environment; you don't need to administer one.
  • Experience with a workflow orchestrator such as Airflow, Prefect, or Dagster.
  • Solid SQL and experience with at least one NoSQL or caching layer (Redis, MongoDB, DynamoDB, or similar).
  • Sound CS fundamentals - data structures, algorithms, and the judgment to reason about performance and correctness in a distributed setting.
  • Strong written communication and the ability to work asynchronously across U.S. time zones. We're fully remote; writing clearly is part of the job.
Nice to Have

Genuinely optional. We don't expect any one candidate to have all of these, and we're prepared to teach them.
  • Databricks or Delta Lake specifically
  • Flink, or other stream-processing frameworks beyond Kafka
  • Vector databases (Pinecone, Qdrant, Milvus, pgvector) and hands-on RAG or embedding pipeline work
  • Elasticsearch or OpenSearch
  • Experience parsing messy, unstructured, or multilingual text at scale
  • Deep database performance tuning, or experience with distributed consensus systems
  • Bachelor's degree in Computer Science, Engineering, or a related field


Department Engineering Locations San Francisco, DC, NY Remote status Fully Remote Yearly salary $120,000 - $140,000 Employment type Full-time

Similar Jobs

More Information Technology Jobs

Find similar Data Engineer, AI & Distributed Systems jobs: