Tata Consultancy Services

Data engineer-TACO

Tata Consultancy Services$120K — $160K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in data engineering or related field
  • Strong Python skills with a focus on production application coding
  • Proficiency in Apache Spark and PySpark for distributed data transformations
  • Experience with Apache Airflow for managing workflows
  • Solid understanding of SQL and relational database modeling
  • Familiarity with AWS S3 for data storage
  • Ability to troubleshoot and resolve data pipeline issues effectively

Responsibilities

  • Develop containerized applications using Python and PySpark to process data from AWS S3
  • Implement Anaplan integrations, including authentication and data reconciliation
  • Author and maintain Airflow DAGs and custom operators for orchestration
  • Manage data storage solutions like AWS S3 and PostgreSQL with schema migrations
  • Build and deploy Docker images for Kubernetes workloads
  • Provide production support by diagnosing Spark performance issues and schema mismatches
  • Create and maintain documentation including runbooks and data lineage

Benefits

  • Professional development opportunities
  • Work with a modern tech stack including AWS and Databricks
  • Collaborative team environment focused on best practices
  • Flexible work arrangements
Full Job Description
Data engineer - TACO | Starship Integrations

Required Skills
• Python production application code (modular, config-driven, unit tested).
• Apache Spark / PySpark building and tuning distributed transformations, including diagnosing memory and performance problems.
• Apache Airflow DAG authoring, scheduling and dependency design, debugging failed runs.
• SQL and relational modeling including schema migrations and data reconciliation.
• AWS S3.
• Production support troubleshooting data pipeline failures and root-cause analysis.

Preferred Skills
• Anaplan integration REST/bulk API, Anaplan Connect, or CloudWorks.
• PostgreSQL / RDS and Alembic migrations.
• Docker, Kubernetes, and CI/CD pipelines (Harness or similar).
• Delta Lake / Databricks.

Responsibilities

Integration applications- Containerized Python/PySpark applications that read source data from S3, apply territory and effective-dating business logic, and publish results downstream.

Anaplan integration- Anaplan authentication and bulk import/export APIs, file formatting and upload, and reconciliation of loaded data against source.

Airflow orchestration- Author and maintain DAGs, custom operators, and environment-specific deployment configuration across environments.

Data stores- AWS S3, Delta/Parquet datasets, and PostgreSQL RDS with schema migrations.

CI/CD & runtime- Docker image builds, Kubernetes workloads, and multi-environment promotion pipelines.

Production support- Respond to pipeline failures Spark memory and performance issues, schema drift, and reconciliation mismatches.

Documentation- Runbooks and lineage documentation.

Salary Range- $120,000-$160,000 a year

About Tata Consultancy Services

Tata Consultancy Services (TCS) is an Indian multinational information technology (IT) services and consulting company, headquartered in Mumbai, Maharashtra, India. It is a subsidiary of Tata Group and operates in 149 locations across 46 countries. TCS is the largest Indian company by market capitalization and is ranked 11th on the Forbes Global 2000 list of the world's biggest public companies. TCS is also the second-largest IT services company in the world by revenue and the largest employer of women in India. The company provides services in areas including IT, consulting, and business solutions.
Learn more about Tata Consultancy Services
Size
469,261 employees
Industry

Similar Jobs

More Jobs at Tata Consultancy Services

More Information Technology Jobs

Find similar Data engineer-TACO jobs: