Data Engineer, Forward Deployed

Applied Computing

$110K — $130K *
Information Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • 5-7 years of experience in data engineering or related roles.
  • Deep expertise in PostgreSQL with a focus on performance optimization.
  • Strong proficiency in Python, especially for data processing and scripting.
  • Hands-on experience with AWS services relevant to data pipelines.
  • Familiarity with time-series industrial data and its management.
  • Experience with Databricks and PySpark for data processing tasks.
  • Bonus: Experience in oil and gas or energy sectors.

Responsibilities

  • Ingest data from diverse industrial sources and ensure it is contextualized for AI models.
  • Build real-time and batch data pipelines for a Lakehouse architecture.
  • Manage synchronization between various data sources and AWS storage solutions.
  • Implement and monitor schema changes and data consistency over time.
  • Prepare and maintain data pipelines for AI model training and inference tasks.
  • Optimize database queries and data pipelines for efficiency and performance.
  • Deploy and manage data solutions in AWS Kubernetes with persistent storage.

Benefits

  • Hybrid work model with a mix of remote and in-office requirements.
  • Opportunity to work at the intersection of AI and industrial data engineering.
  • Access to advanced technologies such as Databricks, EKS, and more.
  • Professional development and potential for growth within the energy sector.
  • Engagement in meaningful projects that have a real-world impact on industries.
Full Job Description
The Role

As our Data Engineer, you'll architect and maintain pipelines that make high-frequency time-series, lab, and historian data into a scalable Lakehouse architecture, usable for both deep learning models and real-time LLMs. You'll be working across AWS (EKS, S3, EBS, KMS, CloudWatch) and Databricks/PySpark, ensuring data is contextualised, synchronised, and optimised for both deep learning models and real-time LLM workloads.

This isn't a traditional ETL role, you'll be solving problems at the intersection of control systems, industrial data engineering, and AI enablement.

Technical Requirements
  • Deep expertise in PostgreSQL (partitioning, indexing, query optimisation, storage design).
  • Strong proficiency in Python for data processing, scripting, and pipeline orchestration.
  • Hands-on experience with AWS (EKS, S3, EBS, IAM, KMS, CloudWatch, etc.)for secure and scalable data pipelines.
  • Proven ability to work with Databricks and PySpark for large-scale distributed data processing.
  • Familiarity with time-series industrial data (control systems, DCS/SCADA logs, process historians).
  • Experience in unstructured data sync and management within hybrid cloud/on-prem environments.
  • Bonus: Experience working as a data engineer in oil and gas or energy environments
  • Bonus: Knowledge of streaming frameworks (Kafka, Flink, Spark Streaming) or MLOps stacks for data versioning and lineage.

Core Responsibilities

1. Ingest & Contextualise Data
  • Ingest from OPC UA servers, process historians, IoT sensors, LIMS systems, alarms/events, and P&IDs.
  • Map signals to their physical processes (tags, units, hierarchies) for interpretability in AI pipelines.

2. Data Movement & Accessibility
  • Build pipelines that handle real-time streaming and batch ingestion into the Lakehouse.
  • Manage synchronisation between historian archives, unstructured files, and AWS storage (S3/EBS).
  • Orchestrate Databricks Lakeflow/Connectors for integrating data into Lakebase/Lakehouse.
  • Handle secure, high-throughput transfers between historian archives and sandbox/live environments.

3. Change Tracking & Integrity
  • Detect and manage schema changes, signal drift, and inconsistencies acrosstime.
  • Implement lineage and audit trails across Spark/Databricks and AWS pipelines.


4. Data Preparation for AI
  • Build and maintaindual pipelines:
    • Training→ large-scale historical data prep for time-series + LLM training.
    • Inference→ low-latency, real-time pipelines for anomaly detection, optimisation, and LLM search.
  • Support heterogeneous AI workloads (time-series forecasting and retrieval-augmented LLMs).

5. Database Performance & Optimisation
  • Tune PostgreSQLand sparkfor high-throughput time-series workloads (partitioning, indexing, query optimisation).
  • Optimise pipelines for both fast analytical queries and high-efficiency model training.
  • Deploy and manage data pipelines in AWS EKS (Kubernetes) with persisten tEBS-backed storage.


What Success Looks Like
  • Live data streams are contextualised,queryable, and AI-ready.
  • Schema changes and signal drift are detected and handled without breaking downstream workflows.
  • Training and inference pipelines run smoothly in parallel, optimised for scale and latency.


Department Commercial Role Forward Deployed Locations Houston Remote status Hybrid

Similar Jobs

More Jobs at Applied Computing

More Information Technology Jobs

Find similar Data Engineer, Forward Deployed jobs: