Data Engineering Scientist

Maya HTT

$110K — $130K *
US-AnywhereRemote in Ontario, CA
Enterprise Technology
Less than 5 years of experience
Job Overview by Ladders

Qualifications

  • Bachelor's, Master's, or PhD degree in a relevant field
  • Strong background in data engineering, mining, and statistical modeling
  • Proficiency with data engineering toolkits, especially OSIsoft Aveva PI
  • Familiarity with geo-spatial tools like Esri ArcGIS is advantageous
  • Hands-on experience with time-series databases, ideally from industrial sensors
  • Excellent scripting abilities in Python, SQL, or PowerShell
  • Experience with cloud solutions like Aveva and AWS

Responsibilities

  • Collaborate with data scientists to build and maintain efficient data pipelines
  • Utilize Aveva PI AF, EF, and related tools for data management
  • Ensure agile, proper execution of data validation and merging
  • Define and monitor business-relevant data quality metrics
  • Identify dataset biases and ensure data cleanliness
  • Experiment with, build, and optimize data pipelines and methodologies
  • Use data mining techniques to identify patterns for data scientists and experts
  • Enhance data collection strategies for engineering and manufacturing systems
  • Work with other specialists to guarantee data integrity for ML-Ops
  • Present data findings clearly to stakeholders
  • Support the continuous performance tracking of ML/AI models
  • Integrate time-series data with geo-spatial tools for better visibility

Benefits

  • Collaborative work environment with a focus on innovation
  • Opportunity to work with cutting-edge AI and ML technologies
  • Access to a team of experts in engineering and manufacturing
  • Agile work culture allowing for flexibility and creativity
  • Potential for professional growth and knowledge expansion in high-demand skills
Full Job Description
Role Summary:

As a Data Engineering Scientist, you'll work closely with a team of data scientists and engineering/manufacturing/operations subject matter experts, mostly with data coming from operational technologies (OT), historians, industrial sensor data, video streams data, and audio stream data. You will also work with engineering and manufacturing experts in order to solve industrial operations business use cases by creating data pipelines to feed ML-Ops systems running machine learning, deep learning and advanced analytics in general.

You will be called upon to align the data available, how to access it securely, transform it, merge and fuse it with other data streams, in order to pre-process the right data and at the right frequency, and make that data reliably useable by the right AI technology(ies). Your focus as part of the team is to select the appropriate data pipelines to best solve the engineering and manufacturing challenges posed collaboratively by our clients and our own Maya engineering and manufacturing experts. You will collaborate with other team members and architect the right data pipelines which will help create a practical AI-Ops (or ML-Ops) solutions for effective decision making by Maya HTT's clients. As is often the case in newer engineering and manufacturing applications leveraging AI technologies, you will work in an agile fashion to build data pipelines and design experiments to extract the value within the data provided. Given Maya HTT's focus in engineering and manufacturing AI applications, excellent understanding of time-series from industrial sensors is a big plus, and video processing and audio data processing is a plus.

What to expect as your main responsibilities:
  • Collaborating with data scientists to create the right data pipelines
  • Working with Aveva PI AF, EF and other PI tools
  • Ensuring the data engineering is done properly while remaining agile in the gradual data validation, data merging, and in general produce the right data aligned to the business use case to solve
  • Defining the data quality metrics to be tracked from a business perspective, and metrics to be optimized on
  • Checking data cleanliness and identifying dataset biases whenever relevant
  • Experimenting, building, and optimizing selected data pipeline and methods
  • Leveraging data mining to uncover interesting patterns or correlation using state-of-the-art methods to help data scientists and subject matter experts
  • Augmenting datasets either algorithmically or using third party sources of information when needed
  • Enhancing data collection procedures to include information that is relevant for building better engineering and manufacturing automation & optimization systems
  • Interacting with data engineering specialists to ensure data processing, cleansing, and verifying the integrity of data used for the ML-Ops is reliably done and available during 24/7 operations
  • Presenting data mining and early data finding results in a clear manner
  • Collaborating with 24/7/365 ML-Ops team and ensuring the tracking of data feeding the ML/AI model performance over time
  • Integrating time-series data into geo-spatial tools for increased data visibility in industry


Minimum Requirements:
  • A Bachelor, Master degree or PhD
  • Excellent understanding of data engineering, data mining, and statistical modelling
  • Experience with common data engineering toolkits such as ex-OSIsoft Aveva PI (a must have), Insights Hub, Azure IoT, AWS IIoT/Sitewise
  • Experience with geo-spatial toolkits such as Esri ArcGIS is a plus
  • Experience with timeseries database, especially if you have hands-on experience with industrial sensors data
  • Excellent scripting skills (typically python, SQL, PowerShell)
  • Experience with cloud hosting and solutions (Aveva, AWS, etc)
  • Good data story communication skills is a big plus
  • Experience with data visualisation tools
  • Experience using query languages such as SQL
  • Experience with NoSQL databases
  • Good teamwork skills

Similar Jobs

More Enterprise Technology Jobs

Find similar Data Engineering Scientist jobs: