Full Job Description
Machine Learning Engineer II
As a Machine Learning Engineer II, you will design, build next-generation predictive and prescriptive maintenance systems utilizing Azure Machine Learning Studio from end to end. Drawing on industrial sensor data and machine PLC , you will develop cutting-edge models that detect failure signatures before they occur and prescribe optimized corrective actions.
You will own the end-to-end industrial ML lifecycle. You will design, train, and optimize supervised and unsupervised architectures to accurately predict equipment Remaining Useful Life (RUL), detect complex anomalies, and deploy prescriptive Agentic AI decision workflows. Once validated, you will deploy these models to low-latency cloud and edge endpoints, seamlessly integrating predictions with plant dashboards, End points Applications and CMMS workflows. Finally, you will establish automated MLOps pipelines in Azure ML Studio to continuously monitor data drift and trigger zero-downtime model retraining as physical factory environments evolve.
This is a high-impact & cross-functional engineering role requiring strong technical depth, system-level thinking, and the ability to communicate complex solutions effectively to diverse audiences including operations (Manufacturing plants), IT, systems Engineering, and executive leadership.
Key Responsibilities
1. Advanced ML Modeling & Algorithmic
Supervised & Unsupervised Learning: Build robust classifiers for fault diagnosis and regression models for Remaining Useful Life (RUL) estimation. Expertly handle highly imbalanced datasets where failure labels are rare.
Agentic AI & Prescriptive Systems: Develop multi-agent workflows that reason over asset health data, parse digital manuals via RAG (Retrieval-Augmented Generation), interact with operational APIs, and generate automated outputs.
Utilizing XGBoost, Random Forests, LSTMs, and Autoencoders—to process sensor streams and PLC data for predictive maintenance and real-time anomaly detection. leverage techniques like Isolation Forests, One-Class SVMs, Dynamic Time Warping, and PCA to build scalable models that monitor asset health, classify process quality, and drive automated decision-making.
2. Production-Grade MLOps & Infrastructure
Robust Data Engineering: Standardize, clean, and enrich raw, unstructured, or missing sensor telemetry and PLC tag data.
Scalable ML Pipelines: Build and maintain scalable, reproducible training and inference pipelines (using MLflow, Kubeflow, or Azure Machine Learning).
Edge & Cloud Deployment: Deploy models across hybrid environments, optimizing for cloud (Azure) as well as low-latency.
Distributed Compute Tuning: Optimize model training and throughput, leveraging GPU-accelerated training and efficient serialization for massive datasets.
3. Systems Integration & Cross-Functional Impact
High-Fidelity Code: Deliver highly optimized, production-grade, modular software in Python and C++ accompanied by strict unit testing, and clean documentation.
Technical Communication: Bridge the gap between data science and physical operations. Clearly articulate complex ML mechanics, decision boundaries, and model limitations to plant managers, IT directors, and executive leadership.
Qualifications & Deep Technical Requirements
Technical Skills (Must-Haves):
Frameworks & Libraries: Deep expertise in PyTorch or TensorFlow, alongside standard data science libraries (Scikit-Learn, NumPy, Pandas, SciPy).
Production Programming: Exceptional software development skills in Python (writing optimized, vectorized code) ,Java Script, C/C++ & R
Modern MLOps & Cloud: Hands-on experience with containerization (Docker/Kubernetes), distributed processing (PySpark/Databricks), and cloud architectures, ideally Microsoft Azure.
Data Handling: Mastery of SQL, NoSQL, and time-series databases (e.g., InfluxDB, TimescaleDB) containing millions of streaming data points.
This position is estimated to travel 10-30%
Please note this job description is not a full list of activities, duties or responsibilities required of the employee for this job. Duties, responsibilities, and activities may change at any time with or without prior notice.
This position embodies the values of Niagara’s LIFE competency model, focusing on the following key drivers of success:
Lead Like an Owner
Manages a safe working environment, accurately documents safety-related training, and effectively communicates safety incidents
Provides strategic input and oversight to departmental projects
Makes data-driven decisions and develops sustainable solutions
Skilled in reducing costs and managing timelines while prioritizing long-run impact over short-term wins
Makes decisions by putting overall company success first before department/individual success
Leads/facilitates discussions to get positive outcomes for the customer
Makes strategic decisions that prioritize the needs of the customer over departmental/individual goals
InnovACT
Continuously evaluates existing programs and processes, and develops new initiatives to increase efficiency and reduce waste
Creates, monitors, and responds to departmental performance metrics to drive continuous improvement
Champions responsible adoption of Agentic AI and intelligent automation to improve reliability, speed, decision quality, and waste reduction while maintaining safety and governance.
Communicates a clear vision, organizes resources effectively, and adjusts the strategy as needed when managing change
Find a Way
Demonstrates ability to think analytically and synthesize complex information
Effectively delegates technical tasks to subordinates
Works effectively with departments, vendors, and customers to achieve organizational success
Identifies opportunities for collaboration in strategic ways
Empowered to be Great
Makes hiring decisions primarily based on culture fit and attitude, and secondarily based on technical expertise
Engages in long-term talent planning
Provides opportunities for the development of all direct reports
Understands, identifies, and addresses conflict within own team and between teams
Work Experience/KSA’s
Required:
Education: Bachelor’s degree in Computer Science, Data Science, Electrical/Mechanical Engineering, Mathematics, or related quantitative field (Master’s or PhD with an ML focus is highly preferred).
Strong proficiency in Python and modern machine learning frameworks such as PyTorch and/or TensorFlow
Experience working across multiple modalities, with expertise in one or more of:
Natural Language Processing: LLMs, text classification, information extraction, retrieval systems, agentic applications, or related areas.
Experience training, fine-tuning, evaluating, and deploying machine learning models in production environments.
Experience designing evaluation methodologies, benchmarking systems, and model performance metrics.
Experience with MLOps tools and practices (Docker, Kubernetes, CI/CD for ML, MLflow, etc.)
Experience with cloud platforms such as Google Cloud Platform (preferred), AWS, or Azure, including ML infrastructure, workflow orchestration, storage, and database services.
Preferred:
Masters or PhD degrees are preferred.
5-7 years - Experience in Python, R, or another programming language
5-7 years - Experience with TensorFlow, PyTorch, scikit-learn, or comparable ML frameworks
5-7 years - Experience in Industrial ML, Automation, Data Science, AI, or related fields
5-7 years - Experience with cloud computing platforms such as AWS, Azure, or GCP
3-5 years - Experience with natural language processing (NLP), LLM applications, prompt engineering, or retrieval-augmented generation (RAG)
3-5 years - Experience leading production Agentic AI, LLM, RAG, or multi-agent orchestration initiatives in industrial, manufacturing, maintenance, reliability, or enterprise operations environments
5-7 years - Experience with Deep Learning, Computer Vision, Reinforcement Learning, or advanced predictive modeling
3-5 years - Experience with ethical, legal, privacy, security, and responsible AI considerations in machine learning and agentic AI systems
3-5 years - Experience with AgentOps/LLMOps practices, including monitoring, evaluation, versioning, safety testing, audit trails, and cost/performance optimization
*Experience may include a combination of work experience and education
Preferred Competencies and Skills
Proficiency in Azure ML Studio and related tools for model development, deployment, and monitoring.
Proficiency in Agentic AI and LLM application development, including prompt engineering, RAG, vector search/embeddings, function/tool calling, and agent workflow orchestration.
Experience with Agentic AI frameworks or platforms such as LangChain, LlamaIndex, Microsoft Semantic Kernel, AutoGen, CrewAI, Azure AI Foundry, OpenAI API, or equivalent.
Ability to design secure AI agent integrations with APIs, databases, CMMS/EAM platforms, cloud services, and industrial data sources while enforcing least-privilege access and approval gates.
Ability to evaluate and monitor AI agent performance using offline and online evaluations, trace logs, quality metrics, guardrails, human feedback, and incident response processes.
Understanding of Responsible AI, privacy, prompt-injection risks, model/tool misuse, auditability, and governance for autonomous or semi-autonomous AI agents.
Proficiency in using query languages such as SQL, Hive, Pig. Etc.
Proficiency in, but not limited to:
Microsoft Office Applications – Word, Excel, PowerPoint, Outlook, Project, Visio, etc.
Proficiency in applied statistical skills, such as distributions, statistical testing, regression, etc.
Basic understanding of data acquisition and processing tools and techniques and developing algorithms on common platforms to generate outputs
Scripting and programming skills such as Python, SQL, JavaScript, C++, C#, or API-based integration for analytics, automation, and AI agent tool development
Basic understanding of PLC/SCADA systems such SIEMENS S7, ALLEN BRADLEY, BnR, Edge data Management, etc.
Understanding of machine learning techniques and algorithms, such as k-NN, Naive Bayes, SVM, decision forests, gradient boosting, neural networks, and LLM-based approaches
Preferred experience with common data science and AI toolkits, such as R, Weka, Python, NumPy, Matplotlib, Pandas, MATLAB, Azure ML, and LLM/agent development libraries
Able to translate data, model outputs, and AI agent recommendations into actionable decisions for senior management
Strong analytical and problem-solving skills
Self-motivated with a proven record of taking initiative
Able to work with minimal supervision
Detail-oriented with excellent oral and written communication skills
Able to execute tasks in a very dynamic and ever-changing environment
Education
Minimum Required:
Bachelor's Degree in Computer Science, Data Science, Artificial Intelligence, Industrial/Automation Engineering, or other related fields or equivalent experience
Preferred:
Master's Degree or PhD in Computer Science, Data Science, Artificial Intelligence, Industrial/Automation Engineering, or related field
Certification/License:
Required: N/A
Preferred: N/A
Typical Compensation Range
Pay Rate Type: Salary
$100,464.14 - $145,673.02 / Yearly
Benefits
Our Total Rewards package is thoughtfully designed to support both you and your family:
Regular full-time team members are offered a comprehensive benefits package, while part-time, intern, and seasonal team members are offered a limited benefits package.
Paid Time Off for holidays, sick time, and vacation time
Paid parental and caregiver leaves
Medical, including virtual care options
Dental
Vision
401(k) with company match
Health Savings Account with company match
Flexible Spending Accounts
Expanded mental wellbeing benefits including free counseling sessions for all team members and household family members
Family Building Benefits including enhanced fertility