Data Engineer
PRIMARY PURPOSE: The Gen AI Engineer within the Transformation Office serves as the hands-on architect of the enterprise data supply chain for the organization's most advanced analytics, data science, and AI initiatives. This role performs the critical engineering work required to deliver high-fidelity, production-grade data that powers machine learning models, feature stores, and generative AI applications.
Operating as a "day-one" builder, the Gen AI Engineer designs and delivers data pipelines that bridge legacy on-premise systems-including mainframes, SQL Server, and DB2-with modern cloud platforms such as Snowflake and AWS/Azure AI ecosystems. The role ensures that data is not merely transferred, but deliberately engineered to meet the statistical, performance, and governance requirements of model training, inference, and RAG-based AI systems.
ESSENTIAL FUNCTIONS AND RESPONSIBILITIES- Designs, builds, and maintains resilient ETL/ELT pipelines that ingest data from on-premise systems, AWS services (S3, RDS), and Azure platforms (Blob Storage, Azure SQL), centralizing and curating data for consumption in Snowflake and downstream AI services.
- Develops and maintains feature stores and analytically optimized datasets that support machine learning workflows, ensuring data is clean, versioned, reproducible, and statistically valid for Data Science teams.
- Engineers data pipelines that enable generative AI use cases, including the automated extraction, transformation, chunking, and loading of structured and unstructured data into vector databases across AWS and Azure environments.
- Acts as a Snowflake power user and technical lead, implementing advanced data modeling patterns, Snowpipe automation, and compute and storage optimization to support high-concurrency analytics and AI workloads.
- Executes non-invasive data extraction strategies to unlock mission-critical data from decades-old legacy systems while preserving system stability and avoiding disruption to core business operations.
- Designs and manages complex, cross-platform data workflows using orchestration tools such as Airflow, AWS Step Functions, and Azure Data Factory to ensure reliable, synchronized data movement across the organization's multi-cloud architecture.
- Partners closely with central IT, database administrators, infrastructure, and security teams to resolve connectivity and access challenges-including PrivateLink, IAM, network segmentation, and firewall controls-while securing production approval for new data integrations.
- Implements automated data quality, validation, and observability frameworks to detect data drift, anomalies, and integrity issues that could negatively impact production analytics, machine learning, or AI systems.
- Drives efficiency across the data ecosystem by optimizing storage, compute usage, and query performance in Snowflake, AWS, and Azure, ensuring responsible cost management and measurable ROI for Transformation Office initiatives.
- Operates as a dedicated engineering partner to MLOps, Data Science, and AI teams, rapidly iterating on evolving data requirements and translating experimental use cases into scalable, production-ready data solutions.
ADDITIONAL FUNCTIONS and RESPONSIBILITIES- Performs other duties as assigned.
- Travel as required.
QUALIFICATIONSEducation & LicensingMaster's degree in Computer Science, Data Engineering, or a related field from an accredited college or university preferred.
ExperienceSix (6) years of hands-on data engineering experience, with a track record of building production-grade pipelines for Data Science and AI in multi-cloud environments or equivalent combination of education and experience required.
Skills & Knowledge- Expert-level proficiency in Snowflake architecture, including data sharing, performance tuning, and the integration of Snowflake with external cloud AI services
- Advanced, hands-on knowledge of AWS (S3, Glue, Lambda) and Azure (Data Factory, Synapse) data services
- Mastery of Python, SQL, and PySpark. Deep experience with data orchestration and containerization (Docker)
- Proven ability to interface with "old world" tech (on-premise SQL, Mainframe extracts, flat files) and transform it for modern cloud consumption
- A strong understanding of the specific data needs for Machine Learning (feature engineering) and Generative AI (vectorization and embedding pipelines)
- A "get-it-done" attitude, capable of navigating enterprise bureaucracy and technical debt to ship code at the speed required by a Transformation Office
- Ability to work in a team environment
- Ability to meet or exceed Performance Competencies
WORK ENVIRONMENTWhen applicable and appropriate, consideration will be given to reasonable accommodations.
Mental: Clear and conceptual thinking ability; excellent judgment, troubleshooting, problem solving, analysis, and discretion; ability to handle work-related stress; ability to handle multiple priorities simultaneously; and ability to meet deadlines
Physical: Computer keyboarding, travel as required
Auditory/Visual: Hearing, vision and talking
The statements contained in this document are intended to describe the general nature and level of work being performed by a colleague assigned to this description. They are not intended to constitute a comprehensive list of functions, duties, or local variances. Management retains the discretion to add or to change the duties of the position at any time.