10+ years of software engineering experience, including 2 years in a technical lead role
Expertise in Python with strong fundamentals in testing, code reviews, and API design
Experience with data products using Snowflake, DBT, and Airflow
Proven track record in delivering classical ML and GenAI systems with MLOps practices
Deep knowledge of AWS services for building data and AI platforms
Experience in Infrastructure as Code (IaC) using Terraform or AWS CDK
Familiarity with AI-assisted software development workflows and structuring tasks for AI coding agents
Strong communication skills for engaging with technical and non-technical stakeholders
Responsibilities
Design and build reliable batch and streaming data pipelines
Establish data modeling, quality, lineage, and cataloging practices
Build feature pipelines and curated datasets for analytics and ML use cases
Productionize classical ML models following MLOps practices
Develop GenAI applications for enterprise document retrieval and evaluation
Implement guardrails and observability for LLM-based systems
Define evaluation approaches for measurable AI quality and regression visibility
Create self-service tooling and CI/CD pipelines for consistent product delivery
Own infrastructure as code and reliability practices for the platform
Embed security and compliance requirements into the platform design
Mentor junior engineers in design and implementation best practices
Benefits
Opportunity to work in a cutting-edge AI and analytics environment
Engagement with advanced technologies in the oil and gas sector
Chance to shape the architecture and engineering standards of a modern data platform
Mentorship opportunities for professional growth
Collaboration with cross-functional teams and stakeholders
Full Job Description
We are currently seeking an experienced Data Engineer to join the AI and Advanced Analytics department. As part of the AI Engineering team, the Lead Data Engineer will architect, design, and implement a cloud based modern enterprise Data & AI platform to support AI and analytic use cases for the midstream oil and gas business units. This individual will shape the architecture, set engineering standards, and write production code across data products, AI models, and platform infrastructure.
Responsibilities include, but are not limited to:
Design and build reliable batch and streaming pipelines, including ingestion of industrial time-series and operational data alongside unstructured document sources
Establish data modeling, quality, lineage, and cataloging practices for data product development
Build feature pipelines and curated datasets that serve analytics, ML, and GenAI use cases
Productionize classical ML models following MLOps practices
Build GenAI applications such as retrieval-augmented generation over enterprise documents, including chunking, embeddings, vector search, prompt management, and evaluation
Implement guardrails, observability, and cost controls for LLM-based systems
Define evaluation approaches that make AI quality measurable and regressions visible
Build self-service tooling, reusable templates, and CI/CD pipelines so teams can ship data and AI products consistently on AWS and on-premises environments
Own infrastructure as code, observability, and reliability practices for the platform
Embed security, governance, and compliance requirements into the platform by design
Define and implement best practices around AI Assisted Software Development Life Cycle
Mentor junior engineers in design and code reviews and implementation best practices
Qualifications
The successful candidate will meet the following qualifications:
10+ years of software engineering experience, with at least 2 years in a technical lead role
Expert level experience in Python with solid software engineering fundamentals including testing, code reviews, branching policies, version control, design patterns, and API design
Experience building data products leveraging Snowflake, DBT, and Airflow
Experience delivering classical ML and GenAI systems to production, including MLOps practices such as model registries, pipeline automation, and monitoring
Deep AWS experience building data and AI platforms including S3 Tables, Glue, Athena, Lake Formation, SageMaker, Bedrock, Lambda, ECS, RDS, IAM, and VPC
Experience in designing and implementing Infrastructure as Code (IaC) using Terraform, AWS CDK, or CloudFormation along with deployment automation with CI/CD pipelines
Practical experience with AI-assisted software development workflows using tools such as Claude Code or Cursor to design, implement, test, and refactor code
Experience in structuring work for AI coding agents such as writing clear specifications, providing context, breaking tasks down, and reviewing generated output critically
Clear communication skills and the ability to work with both technical and non-technical stakeholders