Data engineer - TACO | Starship Integrations
Required Skills
• Python production application code (modular, config-driven, unit tested).
• Apache Spark / PySpark building and tuning distributed transformations, including diagnosing memory and performance problems.
• Apache Airflow DAG authoring, scheduling and dependency design, debugging failed runs.
• SQL and relational modeling including schema migrations and data reconciliation.
• AWS S3.
• Production support troubleshooting data pipeline failures and root-cause analysis.
Preferred Skills
• Anaplan integration REST/bulk API, Anaplan Connect, or CloudWorks.
• PostgreSQL / RDS and Alembic migrations.
• Docker, Kubernetes, and CI/CD pipelines (Harness or similar).
• Delta Lake / Databricks.
Responsibilities
Integration applications- Containerized Python/PySpark applications that read source data from S3, apply territory and effective-dating business logic, and publish results downstream.
Anaplan integration- Anaplan authentication and bulk import/export APIs, file formatting and upload, and reconciliation of loaded data against source.
Airflow orchestration- Author and maintain DAGs, custom operators, and environment-specific deployment configuration across environments.
Data stores- AWS S3, Delta/Parquet datasets, and PostgreSQL RDS with schema migrations.
CI/CD & runtime- Docker image builds, Kubernetes workloads, and multi-environment promotion pipelines.
Production support- Respond to pipeline failures Spark memory and performance issues, schema drift, and reconciliation mismatches.
Documentation- Runbooks and lineage documentation.
Salary Range- $120,000-$160,000 a year