Full Job Description
We are seeking a Senior Data Engineer to join our newly formed centralized analytics team as one of the first Data Engineers on the team. This is a greenfield opportunity to build a data platform from the ground up-making foundational architectural decisions and directly influencing how an entire organization measures success and makes investment decisions. You will design, build, and operate scalable data pipelines that connect product telemetry, usage metrics, and business outcomes into a coherent, unified data ecosystem. Your focus will be squarely on engineering-building robust, scalable infrastructure and data models-while dedicated Business Intelligence Engineers on the team own the reporting, dashboarding, and stakeholder-facing analytics. This is not traditional reporting-you will be building the data backbone that powers intelligent, agent-driven analytics experiences (MCP tools, agentic retrieval systems) enabling stakeholders to intuitively access and consume data within their day-to-day workflows. The data you engineer will inform executive reviews, drive product strategy, and power the next generation of self-service analytics tools used by thousands of AWS field team members.
Key job responsibilities
- Architect and own the end to end data platform strategy for the STT product portfolio, designing scalable ETL/ELT pipelines that ingest product telemetry, usage events, and business outcome data from multiple heterogeneous sources using AWS-native technologies (Redshift, S3, Glue, Lake Formation, Lambda, Athena, MWAA, EMR, Data Zone)
- Define and drive the next generation data architecture for the organization improving scale, quality, and performance while establishing the technical vision and roadmap that aligns data infrastructure investments with business priorities
- Design and implement a centralized data platform serving as the single source of truth for organizational analytics, building and maintaining data models that connect product usage signals to business outcomes (e.g., content effectiveness to field engagement to pipeline progression to revenue impact)
- Lead the development of data infrastructure supporting AI/ML pipelines and agentic systems, including MCP tools and natural-language data access layers, contributing to the evolution from static dashboards toward agentic data systems by building the foundational data layers that AI agents query and reason over
- Establish and enforce data governance best practices including data contracts, lineage tracking, catalog metadata, data quality frameworks with automated monitoring, alerting, and validation to ensure accuracy, consistency, compliance with security and privacy regulations, and trust across the organization
- Build self-service data products with clear SLAs, documentation, and governance that reduce ad-hoc request burden and empower stakeholders to answer their own question, developing and maintaining automation scripts to generate structured datasets with focus on efficiency and scalability
- Partner with and provide technical guidance to Applied Scientists, SDE teams, and data consumers to provide clean, well modeled data for agent evaluation frameworks, retrieval quality measurement, content effectiveness scoring, and capacity simulations
- Improve existing solutions by identifying and driving cross team technical improvements, influencing engineering best practices, and raising the bar on data engineering standards across the organization
- Operate with a high bar for operational excellence owning on call, monitoring pipeline health, proactively resolving data freshness or quality issues before they impact consumers, and mentoring junior engineers on operational rigor
- Provide technical leadership and mentorship to data engineers on the team, setting technical direction, conducting design reviews, and elevating the team's overall engineering capabilities
About the team
You will be joining a high-growth engineering organization at the forefront of applying generative AI and agentic technologies to transform how AWS field teams operate. The centralized analytics team is being built from the ground up-you will be one of the first two Data Engineers on the team, working alongside Business Intelligence Engineers, a Senior BD, an Applied Scientist, and a TPM. You will make foundational architectural decisions that define how the platform will be built, scaled, and operate for years to come. The pace of innovation is high, the problems are ambiguous, and the impact is measured across thousands of field team members and the customers they serve. This role offers the opportunity to shape foundational architecture decisions and influence how an entire organization consumes and acts on data.
BASIC QUALIFICATIONS
- 7+ years of data engineering experience
- Experience with data modeling, warehousing and building ETL pipelines
- Experience with SQL
- Experience in at least one modern scripting or programming language, such as Python, Java, Scala, or NodeJS
- Experience mentoring team members on best practices
PREFERRED QUALIFICATIONS
- Experience with big data technologies such as: Hadoop, Hive, Spark, EMR
- Experience operating large data warehouses
The base salary range for this position is listed below. Your Amazon package will include sign-on payments and restricted stock units (RSUs). Final compensation will be determined based on factors including experience, qualifications, and location. Amazon also offers comprehensive benefits including health insurance (medical, dental, vision, prescription, Basic Life & AD&D insurance and option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage), 401(k) matching, paid time off, and parental leave. Learn more about our benefits at https://amazon.jobs/en/benefits.
USA, NY, New York - 170,000.00 - 230,000.00 USD annually
USA, TX, Austin - 154,600.00 - 209,100.00 USD annually
USA, WA, Seattle - 154,600.00 - 209,100.00 USD annually