Backend Engineer - Data Pipeline

Sunsets HQ Corp.

$130K — $155K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 3+ years of professional software engineering experience with production backend systems
  • Strong understanding of asynchronous or distributed systems, including queues and concurrency
  • Experience in environments where output correctness is independent of service health
  • Ability to define invariants and utilize various verification techniques
  • Experience designing versioned data contracts and safely migrating them
  • Skillful in debugging across system boundaries to improve systems post-incident
  • Fluency with modern AI engineering tools and their application in development

Responsibilities

  • Design and implement backend systems for high-volume, multi-stage data processing
  • Define and maintain authoritative, versioned data contracts and manifests
  • Ensure robust handling of retries, checkpoints, and partial failures
  • Build mechanisms for independent verification instead of relying solely on job success
  • Transform failures into structured recovery paths and tests
  • Provide clear insights into the run state and safety controls for product teams
  • Analyze system behavior and improve correctness and efficiency without compromising security

Benefits

  • Collaborative work environment with cross-functional teams
  • Opportunity to leverage AI in software development
  • Focus on improving system reliability and data integrity
  • Chance to address complex data challenges in sensitive environments
  • Professional growth through tackling high-impact engineering problems
Full Job Description
The Role

We're hiring a backend engineer to make the pipeline that de-identifies sensitive enterprise data correct, replayable, operable, and safe to change.

This is not a conventional data-platform role where a successful job is enough. A pipeline can finish while silently dropping records, duplicating output, applying stale policy, losing lineage, or producing evidence that cannot establish whether a dataset is safe to release. You will own the backend systems and contracts that make those failure modes visible, preventable, and recoverable.

You will work across asynchronous orchestration, batch workers, queues, object storage, databases, many file formats, model-backed stages, deterministic verification, and human review. The role is backend-focused, but the outcome is a product and delivery promise: the team must know what ran, what changed, what remains uncertain, and what can safely happen next.
What You'll Own
  • Design and ship backend systems for multi-stage, high-volume data processing
  • Define authoritative, versioned contracts for manifests, artifacts, lineage, and state transitions
  • Make retries, checkpoints, partial failures, replay, backfills, migrations, and rollbacks safe and understandable
  • Build independent reconciliation and verification instead of treating job success as proof of correct output
  • Turn escaped and recurring failures into fixtures, regression coverage, release gates, and durable recovery paths
  • Expose trustworthy run state and safe controls to the products people use to investigate and release data
  • Diagnose production behavior across code, queues, stores, artifacts, data formats, and deployed versions
  • Improve correctness, throughput, and operating leverage without weakening privacy, security, or release confidence
  • Use AI deeply in development and in bounded verification systems, with explicit evaluation and independent checks
  • Partner with machine learning, applied science, full-stack product, platform, security, and data engineering teammates
What Success Looks Like
  • A high-risk pipeline boundary has an explicit contract, independent reconciliation, replayable coverage, and safe recovery
  • Missing, duplicated, stale, or incompatible work is detected before it becomes a customer delivery
  • Material pipeline state and release decisions are backed by queryable provenance and audit evidence
  • Recurring reruns, manual interventions, and diagnosis or recovery time decline
  • Adjacent engineers can add stages and checks through supported patterns instead of one-off scripts and implicit storage conventions
  • Pipeline changes can be rolled out, backfilled, quarantined, or reversed without delivery heroics
You Might Thrive Here If
  • You have at least three years of professional software engineering experience, including personal ownership of production backend systems
  • You are strong in asynchronous or distributed systems and can reason precisely about queues, concurrency, state, storage, idempotency, partial failure, and recovery
  • You have worked on systems where output could be materially wrong even when every service looked healthy
  • You define invariants and use reconciliation, control totals, diffs, replay, goldens, shadow paths, or independent sources to verify correctness
  • You can design versioned data and artifact contracts and migrate them safely in a live system
  • You debug from evidence across system boundaries and turn incidents into durable system improvements
  • You choose technical work based on operator and customer consequences, not architecture in isolation
  • You use modern AI engineering tools fluently, verify their output, and know when model-backed checks need deterministic guardrails and human review
  • You communicate clearly across product, ML, data, platform, security, and customer-facing teams
This Role May Not Be for You If
  • You want to focus primarily on frontend product development or visual craft
  • You treat pipeline success, uptime, latency, or a green dashboard as sufficient evidence that the output is correct
  • You prefer isolated infrastructure work without responsibility for data and delivery consequences
  • You solve partial failure primarily with retries and manual runbooks
  • You do not want AI tools to be part of your daily engineering workflow
Bonus
  • Experience with large-scale batch processing, workflow orchestration, event-driven systems, or data movement
  • Experience with schema evolution, manifests, lineage, CDC, migrations, reindexing, or backfills
  • Experience in payments, ledgers, reconciliation, claims, fraud, identity, search quality, observability, or another domain with delayed or weak ground truth
  • Experience with sensitive or multi-tenant data, least-privilege systems, auditability, quarantine, and fail-closed release paths
  • Experience combining deterministic checks, synthetic fixtures, offline replay, model-based judges, and human review
  • Experience with Python, AWS, Airflow, Batch, SQS, S3, DynamoDB, PostgreSQL, or comparable systems

Similar Jobs

More Jobs at Sunsets HQ Corp.

  • Backend Engineer - Data Pipeline
    $130K — $155K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Data Scientist
    $120K — $145K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Security Lead
    $150K — $180K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person
  • Design Engineer
    $110K — $130K *
    New York, NY 10025 (New York County)
    Enterprise Technology
    In-Person
  • Machine Learning Engineer
    $130K — $160K *
    New York, NY 10025 (New York County)
    Information Technology
    In-Person

More Information Technology Jobs

Find similar Backend Engineer - Data Pipeline jobs: