As a Data Engineer on this team, you will build and scale the data infrastructure behind this source-of-truth dataset, enabling finance and operations teams to assess network execution and understand cost drivers. This role sits at the center of network topology analytics, where the ability to understand Amazon's complex supply chain network and drive data model, mapping, and dimensional standardization is critical. You will onboard 100+ metrics into our central "metrics-as-a-code" platform enabling self-service reporting and analytics for finance and controllership teams. You will design, implement, and support scalable data solutions and complex data models for Amazon's Operations Finance organization. Your work will integrate heterogeneous data sources, automate ETL workflows, and deliver curated datasets used for reporting, analysis, and ad-hoc requests. You will partner with finance stakeholders, BI Engineers, and product teams to gather requirements, architect pipelines, and ensure that data is reliable, accessible, and trusted for data-driven business decisions.
Key job responsibilities
- Design, build, and maintain scalable ETL pipelines and data solutions using Python, SQL, Spark, and AWS services (S3, Glue, Redshift, Airflow, EMR, Lambda)
- Create coherent logical and physical data models that drive the metadata-driven configuration layer, enabling finance teams to define new metrics and views without code changes
- Collaborate with finance stakeholders, BI Engineers, and partner teams to gather requirements, integrate automated and manual data feeds, and deliver trusted datasets
- Participate in code reviews, design discussions, and team planning while identifying process improvements that increase automation and reduce manual error
- Optimize query performance and resource usage across the data platform, ensuring reliability and operational excellence at scale
A day in the life
You start your morning reviewing pipeline health dashboards and triaging data quality alerts on cost data flowing across transportation segments. Mid-morning, you collaborate with finance analysts to validate cost allocation logic for a carrier segment being onboarded into the platform. After lunch you build and test an ETL module that integrates a partner team's operational data feed, then pair with a teammate on a code review to ensure the solution scales across regions. Your afternoon wraps up with a design session on extending item-level cost granularity to support new reporting dimensions.
BASIC QUALIFICATIONS
- 3+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- Bachelor's degree
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
- Knowledge of distributed systems as it pertains to data storage and computing
PREFERRED QUALIFICATIONS
- Experience with AWS technologies like Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Experience with non-relational databases / data stores (object storage, document or key-value stores, graph databases, column-family databases)
- Experience building large-scale, high-throughput, 24x7 data systems
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, WA, Bellevue - 132,100.00 - 178,800.00 USD annually