DescriptionThe Lead Data Engineer owns Awayday's ingestion and pipeline layer end to end, from each source system through to the landing and staging layers in Snowflake that every domain builds on. It is a foundational mandate. You will build and run the pipelines that bring each source system into the warehouse, direct the offshore engineering team that builds alongside you, and hold the ingestion layer to a standard of reliability the rest of the team can depend on as the company continues to grow.
This is the engineering counterpart to our analysts. Where they own the models and metrics a business domain depends on, you own the data getting there reliably, on schedule, and to a standard the rest of the team can build on top of. The ideal candidate is a seasoned, hands-on data engineer, someone who pairs deep command of SQL, Python, and cloud-warehouse ingestion with the judgment to set standards and the leadership to bring an offshore team along.
The core of the role is building and running the ingestion pipelines, steering the offshore pod's delivery through its tech lead, and owning pipeline reliability. It also underpins our AI initiatives, since the governed data this seat lands and stages is the foundation our agentic assistant retrieves from.
RequirementsKey Responsibilities:Pipeline Ownership and Ingestion- Build and maintain the ingestion pipelines that stage every source system feeding the warehouse, spanning our property-management systems and our market, revenue, operations, and web data.
- Bring net-new source systems online end to end: API and file-drop ingestion, landing and staging layers, schema-drift handling, and the load strategies that keep them production-grade.
- Own pipeline performance and cost, from orchestration efficiency to the cost-efficiency of the ingestion tier.
- Stay accountable for the reliability and health of the ingestion layer, not only for shipping new pipelines.
Engineering Team Leadership- Set sprint direction, priorities, and delivery expectations with the offshore tech lead, who assigns work, runs QC and pull-request sign-off, and manages the pod day to day.
- Maintain the technical bar on changes that touch shared ingestion standards or architecture, with the pod running peer review on the rest.
- Grow the pod toward owned outcomes measured against a scorecard, rather than billed hours.
- Coach and mentor the team's engineers, growing the team's capability and the quality of its output over time.
Reliability and Operations - Direct the offshore pod's response to the production bug queue and pipeline incidents, prioritizing and unblocking as needed.
- Establish and run a SEV1 to SEV3 severity scale, incident response, and a monthly reliability report that separates inherited debt from new defects.
- Uphold the debt-paydown capacity the operating model reserves, so reliability improves as the estate grows.
Standards, Governance, and Platform- Enforce the technical data contract each new source meets at ingestion (schema, types, freshness, and quality) to the standard the function lead sets.
- Keep CI gates on every merge and hold the ingestion coding standards the pod builds to.
- Implement the scoped, time-boxed contractor access model in Snowflake and AWS (roles as code, just-in-time elevation) to the governance standard set with IT Security.
Automation and Leverage- Generalize and automate new-shop onboarding for our property-management systems, so a new source type needs an engineer once and a new shop within an existing type is a config row and a checklist.
- Decouple acquisition-driven brand growth from engineering headcount through repeatable, config-driven onboarding.
AI and RAG Enablement- Contribute to the team's AI and RAG initiatives: build and maintain the governed data and retrieval plumbing our agentic assistant depends on.
Qualifications:- 6+ years in data engineering, with demonstrated ownership of production ingestion and ELT pipelines at scale.
- Expert SQL and strong Python for data engineering.
- Deep hands-on experience with a cloud data warehouse for high-volume landing and staging, Snowflake strongly preferred.
- Workflow orchestration in production, using Apache Airflow, AWS MWAA, or an equivalent.
- API- and file-based ingestion (REST APIs, S3), incremental loads, schema-drift handling, and data-quality testing.
- CI/CD for data pipelines and disciplined code review.
- Experience leading engineers or delivering through a team, setting technical direction and standards others build to.
- Powerful written and verbal communication skills to align an offshore team across time zones and to translate technical tradeoffs for non-technical stakeholders.
- Practical, hands-on experience applying AI tools to real work, including LLM-assisted workflows and retrieval-augmented generation (RAG). You know how to structure prompts and context and ground models on trusted data sources. Deep AI-engineering theory is not required, only fluency putting these tools to work.
Preferred Qualifications:- Infrastructure as code (Terraform) and cloud security or RBAC design.
- AWS data services beyond MWAA, such as S3, Glue, and Lambda.
- Experience building metadata-driven or config-driven pipeline frameworks.
- Openness to an evolving data stack, including evaluating orchestration, ingestion, or warehouse tooling beyond our current set as the platform matures.
- Demonstrated leadership of offshore or distributed delivery teams, setting direction and standards and delivering through a tech lead.
Key Attributes:- Ownership mindset. Takes a system from raw source to reliable production and stays accountable for it.
- Raises problems and improvements unprompted, rather than waiting for direction.
- Sets the technical standard and mentors; holding said standard without becoming a bottleneck.
- Calm and methodical under pressure.
- An innate sense of curiosity and desire to help others.