Glossary ofAI

Industrial AI Glossary · 174 essential terms in applied artificial intelligence

AI Glossary: Essential terms in industrial artificial intelligence · Deduce Data Solutions

ThisAI Glossary is DDS's practical resource for understanding the key terms ofArtificial Intelligence applied to industry and energy: Deep Learning, PINNs, XAI, LLM, RAG, Digital Twin, Edge AI or Predictive Maintenance. Each entry in this AI glossary includes a concise definition and a real plant example. If you want to go beyond this AI glossary, you can also check theofficial definition of Artificial Intelligence on Wikipedia or theIEEE standards portal.

A13 terms

Algorithm

A finite, ordered set of instructions that a system follows to solve a problem or produce a result from input data.

Example: the classification algorithm that assigns each part produced to the category "OK" or "defect" on a production line.

Supervised Learning

A type of machine learning in which the model is trained with labelled examples, i.e. known input-output pairs.

Example: training a model with thousands of images already labelled as "defective part" or "correct part" so it learns to classify them on its own.

Unsupervised Learning

A type of machine learning in which the model discovers patterns in the data without pre-existing labels. Useful for grouping, segmenting or detecting anomalies.

Example: grouping energy consumption from different plants to identify typical operating profiles without knowing in advance what they are.

Reinforcement Learning

An approach in which an agent learns to make decisions through trial and error: it receives rewards for good actions and penalties for bad ones.

Example: a system that learns to optimise the scheduling of an industrial furnace by testing combinations and measuring the resulting energy consumption.

AIoTArtificial Intelligence of Things

The convergence of industrial IoT and artificial intelligence: sensors, PLCs and gateways run AI models to make real-time decisions without always relying on the cloud.

Example: a pump with AIoT detects its own vibration deviation and adjusts its operating regime before an operator sees the data in the SCADA.

Autoencoder

A neural network that learns to compress and reconstruct its own input data. It is mainly used to detect anomalies: if the reconstruction departs from the original, something is off in the process.

Example: an autoencoder trained on a motor's normal vibrations automatically detects incipient bearing faults without needing to pre-label each type of failure.

Anomaly Detection

A set of AI techniques that identify data, events or behaviours that deviate significantly from a process's normal pattern. It underpins asset monitoring and industrial cybersecurity.

Example: an anomaly detection system flags an atypical electricity consumption on a line overnight and reveals a compressed-air leak that no one had noticed.

Agentic AI

A system that uses AI agents to carry out tasks autonomously on the user's behalf, capable of deciding, planning and using tools to achieve complex goals.

Example: An agentic AI agent monitors a plant's KPIs and, when a deviation occurs, opens the work order and schedules the inspection without manual intervention.

Active Learning

A machine learning technique in which the model itself selects the most informative data for a human to label, reducing annotation cost and speeding up training.

Example: on a visual inspection line, the system asks the operator to label only the doubtful images and reaches the same accuracy with 70% less labelling.

Attention

A technique that lets a neural network dynamically weigh which parts of the input are most relevant to each prediction, learning to "pay attention" to the key elements of a sequence. It is the core mechanism of Transformers.

Example: a model analysing sequences of turbine sensor data uses attention to give more weight to the relevant vibration moments and anticipate a fault, ignoring the noise.

AI Governance

A set of policies, controls and responsibilities that ensure AI systems are safe, ethical, traceable and compliant with regulations throughout their entire lifecycle.

Example: An industrial group defines its AI governance framework to document, approve and audit each model before it controls critical plant variables.

Autoregressive Model

A model that predicts the next value in a sequence from its previous values. In time series it takes the AR form, and in generative AI it describes how a language model produces each token conditioned on the ones it has already written.

Example: An autoregressive model on a power grid's hourly demand anticipates next hour's consumption using the history of the last 24 hours.

Activation Function

A function each neuron applies to the weighted sum of its inputs to produce its output, introducing non-linearity into the network. Without it, a neural network could only represent linear transformations; ReLU, sigmoid and GELU are the most commonly used.

Example: A network predicting the outlet temperature of an annealing furnace uses ReLU in its hidden layers to capture the non-linear relationship between gas flow, line speed and sheet thickness.
B8 terms

Backpropagation

An algorithm that lets a neural network adjust its internal parameters by propagating the prediction error backwards.

Example: it is the mathematical engine that lets almost any modern neural network learn from its mistakes during training.

Big Data

A set of technologies and techniques for processing data volumes too large, fast or varied for traditional systems.

Example: the telemetry data generated by a connected industrial plant produces gigabytes daily that won't fit in a spreadsheet.

Bayesian Network

A probabilistic model that represents variables and their causal dependencies as a graph. It allows the probability of a future event to be calculated from what has already been observed.

Example: in a chemical plant, a Bayesian network combines pressure, temperature and flow to estimate the probability of a pump failure in the next 24 hours.

Bias

A systematic error in an AI model that produces unfair or skewed results, usually inherited from unrepresentative training data. Detecting and correcting it is key to reliable decision-making.

Example: a quality control model trained only on parts from one shift develops a bias and over-rejects parts from another shift with different lighting.

Batch Normalization

A technique that normalises each layer's activations in a neural network using the mean and variance of the training batch, stabilising and speeding up learning.

Example: by adding batch normalization, the vision model that detects pipe corrosion converges in half the epochs and better tolerates lighting changes.

BERT (Bidirectional Encoder Representations from Transformers)

A Transformer-based language model, introduced by Google in 2018, that interprets each word by considering both its preceding and following context. It is used as a foundation for classifying, searching or extracting information from text.

Example: a BERT model fine-tuned with maintenance reports automatically classifies each incident by fault type and affected equipment.

Beam Search

A decoding algorithm that, at each step of sequence generation, keeps the k most probable hypotheses (the so-called beam width) instead of keeping only the best one, improving the result's quality compared with greedy search.

Example: An assistant drafting maintenance reports uses beam search to choose the most coherent technical wording instead of the first, most probable option.

Bayesian Optimization

A method for optimising expensive-to-evaluate functions that builds a probabilistic surrogate model of the objective and uses an acquisition function to choose the next point to try, balancing exploration and exploitation.

Example: Tuning the operating parameters of a clinker kiln with very few real trials, because each test consumes hours of production.
C15 terms

Computer Vision

The discipline that enables machines to interpret images and video: detecting objects, reading text, identifying defects or measuring distances.

Example: a camera with computer vision detects a defective weld on the production line at a speed of 2,000 parts per hour.

Convolutional Neural NetworkCNN

A type of neural network specialised in processing images. It uses convolutional layers that automatically detect visual patterns (edges, textures, shapes).

Example: almost any modern industrial visual inspection system uses a CNN as its detection engine.

Causal AI

The branch of AI that looks for cause-and-effect relationships, not just correlations. It allows questions like "what would happen if…?" to be answered and interventions with real impact on the process to be designed.

Example: causal AI identifies that the increase in scrap on a line is due to a specific furnace variable, not the raw material as the traditional dashboard suggested.

Clustering

An unsupervised learning technique that groups data into sets (clusters) based on similarity, without prior labels. It is used to segment, explore data and discover hidden profiles.

Example: a clustering algorithm sorts thousands of SCADA alarms into recurring groups and helps maintenance prioritise the most frequent root causes.

Confusion Matrix

A table summarising a classifier's performance by crossing actual classes with predicted ones: true and false positives and negatives. It underpins metrics such as precision, recall or F1.

Example: the quality model's confusion matrix reveals that almost all its errors are false alarms rather than undetected defects, which guides threshold recalibration.

Concept Drift

A change over time in the relationship between the input variables and what the model predicts, so that the learned logic is no longer valid. Unlike data drift, here the concept itself changes, not just the data distribution.

Example: a model estimating a tool's useful life loses validity when the parts' material changes, because the relationship between wear and breakage is no longer the same.

Cosine Similarity

A measure of similarity between two vectors based on the cosine of the angle they form, ranging from -1 to 1, independent of their magnitude. It underpins comparing embeddings by meaning.

Example: A spare-parts search engine calculates the cosine similarity between a fault description and thousands of historical reports to suggest the most similar procedure.

Chain-of-Thought

A prompting technique that asks a language model to reason step by step, showing the intermediate steps before giving the final answer. It improves performance on multi-step tasks and makes its reasoning more transparent.

Example: An assistant diagnosing a compressor fault reasons with chain-of-thought (pressure, temperature and hours of use) before proposing the most likely cause, leaving each step traceable.

Contrastive Learning

A self-supervised learning method that learns representations by pulling similar examples (positive pairs) closer together in vector space and pushing dissimilar ones (negative pairs) apart, without needing labels.

Example: With contrastive learning, a model distinguishes normal and anomalous turbine states from unlabelled vibration signals, generating the pairs through augmentations of the data itself.

Conformal Prediction

A statistical framework that converts any model's point predictions into sets or intervals with guaranteed coverage, assuming only that the data are exchangeable.

Example: In estimating bearings' remaining useful life, it delivers an interval with 90% coverage instead of a single number.

Curriculum Learning

A training strategy that presents examples to the model in order of increasing difficulty, from simple to complex, to speed up convergence and improve generalisation.

Example: training a furnace control model first with stable regimes and later adding start-ups, shutdowns and transients.

Context Window

The amount of text, measured in tokens, that a language model can hold in mind at once. It determines how much information fits between the query and the response before the model starts to forget.

Example: A maintenance assistant with a 128,000-token context window reads a turbine's full manual and answers by citing the exact section.

Counterfactual Explanation

An explanation that indicates the minimal change in the input variables that would have been enough for the model to give a different output.

Example: a model rejects a batch of paint and the counterfactual explanation shows that with 0.4 points less viscosity it would have passed.

Change Point Detection

A statistical technique that identifies the moments when a time series' distribution changes (mean, variance or pattern), splitting it into segments with distinct properties. It is used to detect regime transitions that are not one-off faults but sustained changes.

Example: In an air compressor, the algorithm detects that the vibration mean shifted to a higher level on 12 March, a date that coincides with a badly aligned bearing replacement.

Condition Monitoring

Continuous tracking of a machine's physical parameters (vibration, temperature, current, oil) to detect changes that indicate a developing fault. It is the data foundation that predictive maintenance models work on.

Example: Vibration sensors on the bearings of a boiler's induced-draft fan send data every second to a model that alerts when the spectral signature deviates from the healthy pattern.
D15 terms

Deep Learning

A branch of machine learning based on neural networks with many layers. Capable of learning hierarchical representations of complex data without manual feature engineering.

Example: Deep Learning is what made possible both language models and modern industrial computer vision systems.

Data Drift

A gradual change in the distribution of the data reaching a model in production compared with the data it was trained on. If not monitored, the model silently loses accuracy.

Example: after changing the sheet-metal supplier, the visual inspection model starts misclassifying correct parts because the textures and reflections have changed.

Decision Tree

A model that splits data into branches according to rules like "if X exceeds a threshold, then…", forming an easy-to-interpret tree-like structure.

Example: in maintenance, a decision tree classifies each alarm as "immediate intervention", "planned review" or "ignore" based on hours of use, temperature and vibration.

Digital Twin

A virtual replica of a physical asset, process or plant, fed in real time with real-world data and enriched with AI models to simulate, predict and optimise.

Example: a furnace's digital twin makes it possible to test new combustion strategies virtually before applying them on the plant floor, saving energy without operational risk.

Diffusion Model

A generative model that learns to create new data starting from random noise and reversing a diffusion process step by step. It underpins many of today's image generators.

Example: starting from a sketch and a text instruction, a diffusion model proposes design variants for a part before moving on to the final CAD.

Dropout

A regularisation technique that randomly deactivates a percentage of neurons at each training step, forcing the network to learn redundant representations and reducing overfitting.

Example: a fault-prediction model with little historical data uses dropout to avoid memorising the training cases and generalise better to new faults.

Data Lakehouse

A data architecture that combines a data lake's flexible, low-cost storage with a data warehouse's reliability, governance and query performance, serving both BI and machine learning on a single platform.

Example: an industrial group unifies its plants' raw telemetry and business reports in a data lakehouse, so AI models and dashboards query the same source.

Data Lineage

A record of a piece of data's journey throughout its lifecycle, from its origin and transformations to its destination, enabling the information's reliability to be audited, traced and validated.

Example: When an efficiency KPI comes out anomalous, data lineage allows tracing back from the dashboard to the source sensor and pinpointing the transformation that introduced the error.

Dataset

A structured, organised collection of data (whether tables, images, signals or text) that serves as the basis for training, validating and evaluating machine learning models. Its quality and size directly determine the model's performance.

Example: A dataset with years of a turbine's vibration and temperature readings makes it possible to train the predictive maintenance model that anticipates its failures.

Digital Thread

A continuous flow of data connecting all of a product or asset's information throughout its lifecycle (design, manufacturing, operation and maintenance) enabling traceability and real-time decisions. It is broader than the digital twin, which replicates a specific asset.

Example: A pump's digital thread links its CAD design, its bill of materials, its maintenance history and its telemetry, so that any change is traced end to end.

Data Augmentation

A technique that artificially expands a training set by applying transformations to existing data (rotations, noise, crops, scaling), increasing its diversity to improve generalisation and reduce overfitting.

Example: Faced with few images of a rare defect, data augmentation generates variants with rotations and lighting changes so the visual inspection model recognises it more reliably.

Digital Shadow

A digital representation of a physical asset with automatic data flow only in the physical→digital direction: it reflects the asset's real state but does not act on it, unlike the digital twin.

Example: a panel that replicates an industrial boiler's temperatures and flow rates in real time without sending operating setpoints back to it.

Domain Adaptation

A set of techniques that adjust a model trained in a source domain so it works in a target domain with a different data distribution and little or no labelling.

Example: adapting a defect-detection model trained on one production line to use it on another plant with different cameras and lighting.

Data Mesh

A decentralised data architecture in which each business domain publishes and maintains its own data as a product, with federated governance and a common self-service platform.

Example: In a group with five plants, each one publishes its process history as a data product and the corporate team cross-references them without copying them into a central store.

Dynamic Time Warping (DTW)

An algorithm that measures the similarity between two time series that may run at different speeds, stretching or compressing the time axis to find the best alignment. It outperforms Euclidean distance when two curves have the same shape but are not synchronised.

Example: Comparing each injection press cycle's pressure curve with a reference cycle using DTW makes it possible to detect anomalous cycles even if the machine has slightly varied its rate.
E6 terms

Embedding

A dense numerical representation of a piece of data (word, image, product, signal) in a vector space where similar elements end up close to one another.

Example: in an industrial semantic search system, embeddings make it possible to find a similar maintenance report even if it was described with different words.

Explainable AIXAI

A set of techniques that let an AI model explain its decisions in a way humans can understand. Essential in industrial and regulated environments.

Example: an XAI system doesn't just say "this furnace will fail tomorrow"; it also explains which variables trigger the alert and what action it recommends. This is the default standard at DDS.

Edge AI

Running artificial intelligence models directly on the device (PLC, industrial camera, gateway or machine) instead of sending data to the cloud. It reduces latency, network cost and dependence on connectivity.

Example: a camera with Edge AI on a packaging line detects defects in milliseconds without sending the video outside the plant, also complying with confidentiality policies.

Ensemble Learning

A technique that combines several models to obtain a better, more reliable prediction than any of them individually, reducing bias and variance.

Example: an ensemble of several models predicts the plant's electricity demand by combining their outputs and achieves a lower, more stable error than a single model.

Early Stopping

A regularisation technique that stops training when the error on the validation set stops improving, preventing overfitting.

Example: Training an electricity demand prediction model stops when the validation error fails to improve for ten epochs.

Epoch

One complete pass of the training algorithm through the entire dataset. If the dataset has 10,000 examples and the batch size is 100, one epoch equals 100 iterations; the number of epochs is a hyperparameter tuned with early stopping.

Example: A model classifying defects in laminated sheet metal is trained for 40 epochs, but validation stops improving at epoch 25, so that checkpoint is kept.
F10 terms

Fine-tuning

Adjusting a pre-trained model with data specific to a particular case. It allows a general model to be adapted to a specific domain without training it from scratch.

Example: starting from a general language model and fine-tuning it with thousands of industrial maintenance reports so it understands the sector's technical vocabulary.

Federated Learning

A distributed training technique in which several plants or devices train a shared model without sharing their raw data: only the model updates are exchanged.

Example: three factories from the same industrial group jointly train a quality model without any of them having to hand over sensitive data to the others.

Foundation Model

A large-scale AI model pre-trained on massive amounts of general data (text, code, images) that is later adapted to specific tasks through fine-tuning or prompting.

Example: models such as GPT, Claude, Gemini or Llama are foundation models that, in industry, are specialised to read technical reports, maintenance manuals or grant application documents.

F1-Score

A metric that summarises a classifier's performance as the harmonic mean of precision and recall. It is especially useful when classes are imbalanced.

Example: when detecting defective parts, which are rare compared with correct ones, the F1-score evaluates the model better than plain accuracy, which would be inflated by the majority of good parts.

Feature Store

A centralised repository that stores and manages the features used by ML models, ensuring consistency between training and inference and their reuse across projects.

Example: Several teams at an industrial company reuse the same vibration variables from a feature store for different predictive maintenance models.

Feature Engineering

The process of transforming raw data into variables (features) that better represent the problem for the model, using techniques such as normalisation, aggregation or creating new indicators. It often delivers more improvement than the algorithm itself.

Example: from an accelerometer's raw signal, RMS, dominant frequency and kurtosis are calculated as features that let the predictive maintenance model better detect bearing faults.

Few-shot Learning

An approach in which a model learns to solve a task from a very small number of labelled examples, useful when data are scarce or expensive to annotate.

Example: With just five images of a rare defect, a few-shot learning model learns to recognise it on the visual inspection line.

Function Calling

A capability that lets a language model invoke external tools or functions: it detects when one is needed and generates a structured call (usually in JSON) whose result it incorporates into its reasoning. It is one of the foundations of agentic AI.

Example: Through function calling, a plant assistant queries the historian for furnace 3's temperature and opens a work order in the CMMS without the operator leaving the chat.

Fuzzy Logic

A logical system that admits intermediate degrees of truth between 0 and 1 instead of strictly binary values. Fuzzy controllers apply "if-then" rules over linguistic variables and convert the result back into a numerical signal.

Example: Regulating an industrial furnace's temperature with rules like "if the temperature is low, open the heating valve slightly".

Fault Detection and Diagnosis (FDD)

A control engineering discipline that monitors a system, decides whether a fault exists and locates its type and origin. It combines model-based methods (comparing readings with what a physical model predicts) with signal- or data-based methods.

Example: An FDD system in an industrial HVAC plant detects a drop in chiller performance and attributes it to a stuck expansion valve, not a refrigerant shortage.
G10 terms

GANGenerative Adversarial Network

An architecture in which two neural networks compete against each other: one generates synthetic data and the other tries to tell whether it is real or fake. The result is increasingly realistic samples.

Example: generating synthetic images of rare quality defects to better train a visual inspection model.

Gradient Descent

An optimisation algorithm that adjusts a model's parameters by following the "descent" of the error. It is the mathematical engine behind training in most Machine Learning models.

Example: each training step of an industrial neural network is, internally, a gradient descent move searching for the best possible prediction.

Generative AI

A family of AI models capable of creating new content (text, code, images, CAD designs or process recipes) from natural-language instructions.

Example: an engineer asks generative AI to draft the start-up procedure for a new line following the plant's internal format and safety standards.

Gradient Boosting

An ensemble learning technique that sequentially combines weak decision trees, where each new tree corrects the previous one's errors; XGBoost is its most widespread implementation.

Example: A refinery uses XGBoost on sensor data to estimate the probability of a pump failure in the next 48 hours.

Guardrails

A set of programmable rules and filters placed between the user and a language model to ensure safe, reliable and compliant responses, blocking improper inputs or outputs.

Example: An AI assistant for operators incorporates guardrails that prevent it from recommending manoeuvres outside the plant's safety limits.

GRU (Gated Recurrent Unit)

A simplified variant of a recurrent neural network that uses two gates (update and reset) to capture temporal dependencies with fewer parameters than an LSTM, training faster and with lower memory usage.

Example: A GRU predicts a pumping station's flow rate from the sequence of recent hours, with a lighter model than an LSTM so it can run on the plant gateway.

Genetic Algorithm

An optimisation method inspired by biological evolution: it evolves a population of candidate solutions through selection, crossover and mutation until converging on a sufficiently good solution.

Example: Optimising a wind farm's maintenance shutdown schedule to minimise cumulative generation loss.

Graph Neural Network (GNN)

A neural network that operates on data structured as graphs (nodes and connections) propagating information between neighbouring nodes through message passing, so that the learned representations capture the system's topology.

Example: Predicting congestion and voltage drops in a distribution grid by modelling substations as nodes and lines as edges.

Grid Search

A hyperparameter tuning method that exhaustively evaluates all combinations of a predefined grid of values, usually with cross-validation.

Example: testing all combinations of depth and number of trees for a Random Forest predicting a plant's electricity consumption.

Gaussian Process

A non-parametric method that treats an unknown function as a probability distribution over functions, so that each prediction comes with its uncertainty.

Example: A Gaussian process fits a compressor's performance curve with few trials and flags at which flow rates the prediction is unreliable.
H3 terms

Hyperparameter

A model configuration parameter that is NOT learned during training but decided beforehand (learning rate, number of layers, batch size).

Example: tuning hyperparameters is a critical part of a data scientist's job and determines whether the model learns well or gets stuck.

Hallucination

When an LLM answers with something that sounds coherent but is factually false or made up. It is one of the main risks of using generative AI in critical environments.

Example: an industrial chatbot without RAG invents a non-existent part reference; connected to the ERP via RAG it returns the real catalogue reference and reduces hallucinations to almost zero.

Hidden Markov Model

A probabilistic model of a system that transitions between states which are not directly observed and are inferred from a sequence of measurable observations.

Example: Inferring which wear regime a centrifugal pump is in from the historical sequence of vibration and electricity consumption.
I3 terms

Inference

The phase in which an already trained model is used to make predictions on new data. Unlike training, inference is fast and runs in production.

Example: your industrial plant's predictive model spends most of its useful life doing inference: predicting consumption, anticipating faults, and so on.

Artificial IntelligenceAI

The discipline that develops systems capable of performing tasks that normally require human intelligence: perception, reasoning, language, decision-making.

Example: in industry, AI optimises consumption, anticipates faults, plans production and proposes setpoints to operators, not replacing human decision-making, but supporting it.

Isolation Forest

An anomaly-detection algorithm that builds trees with random partitions: observations that end up isolated with few splits receive a high anomaly score. It does not require labelled data.

Example: Flagging atypical flow and pressure readings in an industrial water network without a labelled fault history.
K3 terms

Knowledge Graph

A data structure that represents entities (equipment, products, people, events) and the relationships between them as a graph, enabling complex reasoning about the domain.

Example: in a refinery, a knowledge graph connects each pump with its incident history, spare parts, drawings and trained technicians, speeding up fault resolution.

Kalman Filter

A recursive algorithm that estimates a dynamic system's state from noisy, incomplete measurements, combining a prediction phase and a correction phase to achieve more accurate estimates.

Example: In a gas turbine, a Kalman filter fuses readings from several vibration and temperature sensors to estimate the shaft's real state.

Knowledge Distillation

A technique that transfers knowledge from a large, accurate model (the "teacher") to a smaller, faster one (the "student"), retaining most of the performance at a fraction of the computational cost.

Example: a distilled visual inspection model runs on the line's Edge camera with almost the same accuracy as the large server-side model.
L8 terms

LLMLarge Language Model

An AI model trained on large amounts of text that can generate, summarise, translate or answer in natural language. GPT, Claude, Llama or Mistral are well-known examples.

Example: an industrial LLM connected to your plant data can answer natural-language questions about historical data, maintenance reports or consumption.

LSTM (Long Short-Term Memory)

A recurrent neural network variant with memory cells and gates (input, forget and output) that captures long-term dependencies and mitigates the vanishing gradient problem.

Example: An LSTM network predicts an industrial plant's daily electricity demand from months of previous consumption.

LoRALow-Rank Adaptation

An efficient fine-tuning technique that freezes the original model's weights and trains only a few small added low-rank matrices, drastically reducing the memory and cost of adapting an LLM.

Example: an engineering team adapts an LLM to its maintenance reports' vocabulary using LoRA on a single GPU, without retraining the base model's billions of parameters.

Linear Programming

A mathematical optimisation method that seeks the best value of a linear objective function subject to linear constraints. It is a classic tool for optimally allocating limited resources.

Example: a plant uses linear programming to distribute the load among several boilers and minimise energy cost while meeting demand and each unit's limits.

LiDAR (Light Detection and Ranging)

A remote-sensing technology that measures distances using laser pulses and generates three-dimensional representations of the environment.

Example: A drone with LiDAR inspects transmission lines and detects vegetation that is too close to the conductors.

Learning Rate

A hyperparameter that sets how much a model's weights change at each gradient descent step. If it is too low, training takes too long; if it is too high, the model oscillates and never converges.

Example: When training an electricity consumption predictor for an extrusion line, lowering the learning rate from 0.01 to 0.001 stabilised the loss, which had previously bounced between epochs.

Load Forecasting

A prediction of the amount of electricity (power in kW or energy in kWh) that will be needed at a given time and place, from hours to years ahead. Grid operators and plants use it to plan generation, purchases and start-ups.

Example: A model trained on consumption history, production schedule and weather forecast anticipates a steel plant's hourly demand to buy energy on the day-ahead market.

Loss Function

A function that quantifies the difference between the model's predictions and the actual values; training consists of adjusting the parameters to minimise it. Mean squared error is common in regression and cross-entropy in classification.

Example: For a model estimating moisture content at a dryer's outlet, mean absolute error is chosen as the loss because it penalises sensor outliers less.
M9 terms

Machine Learning

The branch of AI in which machines learn patterns from data, without being explicitly programmed for each case.

Example: instead of coding rules like "if temperature rises 5°C, alert", an ML system automatically learns which variable patterns precede a fault.

MLOps

A set of practices and tools for deploying, monitoring and maintaining Machine Learning models in production reliably and at scale. The equivalent of DevOps for AI.

Example: MLOps includes retraining pipelines, model drift monitoring and roll-back if a new version worsens results.

Multimodal AI

AI models capable of processing and integrating several types of data at once (text, image, audio, video or signals) to achieve a more complete understanding than a single data type would give.

Example: a multimodal assistant combines a photo of a fault, a vibration reading and the equipment's text history to suggest the most likely cause.

Model Card

A standardised document accompanying an AI model that describes its purpose, training data, metrics, limitations and intended uses, facilitating transparency and auditing.

Example: before deploying the predictive maintenance model, the team publishes its internal model card so quality and safety can validate its usage limits.

Mixture of Experts

A neural network architecture that replaces dense layers with a set of specialised sub-networks ("experts") and a routing network that activates only the most suitable ones for each input, scaling model size at much lower computational cost.

Example: An industrial LLM with mixture of experts routes electrical maintenance queries and chemical process queries to different experts, giving specialised answers without activating the whole model.

Model Registry

A centralised repository that manages the lifecycle of machine learning models: it versions, tags, documents lineage and controls each model's promotion between development, testing and production.

Example: Before a predictive maintenance model controls plant variables, the team saves it in the model registry with its version and metrics, and can roll back to an earlier version if it underperforms.

Model Predictive Control (MPC)

An advanced control technique that, at each instant, uses a dynamic model of the process to predict its future behaviour and solve a constrained optimisation over a moving horizon, applying only the first computed action and repeating the calculation at the next step.

Example: Maintaining product quality in a distillation column while respecting temperature and pressure limits and minimising steam consumption.

Monte Carlo Simulation

A numerical method that estimates results and probabilities by running thousands of simulations with input variables randomly sampled according to their distribution.

Example: A power utility simulates thousands of wind and demand scenarios to estimate the probability of a generation shortfall the following month.

Multi-Task Learning

An approach in which a single model learns several related tasks at once, sharing internal representations that act as regularisation and improve each task's performance.

Example: a single network that simultaneously predicts an extrusion line's product quality and energy consumption.
N3 terms

Neural Network

A computational model inspired by how the brain works, made up of layers of artificial neurons that transform inputs into outputs.

Example: it is Deep Learning's basic building block. Modern neural networks can have millions of neurons distributed across dozens of layers.

NLPNatural Language Processing

The branch of AI concerned with processing and understanding human language: comprehension, generation, translation, text classification.

Example: NLP is what lets an industrial assistant understand the question "how much energy did furnace 3 consume yesterday?" and answer with the exact figure.

Named Entity Recognition (NER)

A natural language processing task that locates entities such as people, organisations, places, dates or quantities in a text and classifies them by category.

Example: On freely handwritten fault reports, NER extracts the equipment, the spare-part code and the date to feed the CMMS history.
O7 terms

Overfitting

A phenomenon in which a model memorises the training data instead of learning generalisable patterns. It works very well on the known and very poorly on the new.

Example: a model with overfitting predicts the last 3 years' data perfectly but fails when an unusual raw-material batch comes in.

OCR

A technology that converts text contained in scanned images or PDFs into editable, software-processable digital text, the basis of industrial document automation.

Example: an OCR automatically extracts batches, dates and references from paper delivery notes and loads them into the ERP without human intervention, eliminating transcription errors.

Object Detection

A computer vision task that locates the objects present in an image and classifies each one, drawing a box around its position. It combines localisation and classification in a single pass.

Example: an object detection system identifies and locates each operator's PPE (helmet, gloves, vest) on plant cameras to verify safety compliance.

One-Hot Encoding

A technique that converts a categorical variable into a numerical vector of N elements, where only the position corresponding to the category is 1 and the rest are 0, allowing models to process non-numerical data.

Example: Representing the type of material processed in each batch as input for a model that predicts the line's energy consumption.

Optical Flow

The apparent motion pattern of pixels between consecutive video frames, represented as a field of displacement vectors.

Example: Optical flow on a conveyor belt camera measures the material's real speed and detects belt slippage.

Observability

A property of a system indicating whether its internal state can be deduced from the measurements available over a finite time interval. If a system is not observable, no estimator will be able to reconstruct those variables.

Example: before installing a Kalman filter on a reactor, it is checked whether the temperature and pressure probes are enough to deduce the internal concentration that no one measures.

OPC UA (Open Platform Communications Unified Architecture)

A vendor-independent industrial communication standard (IEC 62541) that carries plant data along with its semantic description and built-in security.

Example: an energy consumption model reads meters from several packaging lines via OPC UA without needing a different driver for each PLC.
P13 terms

Physics-Informed Neural NetworksPINNs

Neural networks that incorporate known physical equations of the process during training. They combine real data with physical laws for more reliable, defensible predictions.

Example: in steel plants, a PINN respects the furnace's heat-transfer laws in addition to learning from historical data, producing predictions that engineering can technically defend.

Prompt

An instruction or question given to a generative AI model (typically an LLM) to obtain a specific response. The prompt's quality directly affects the response's quality.

Example: prompt engineering is the discipline of writing good instructions to extract maximum value from a language model.

Predictive Maintenance

A maintenance strategy based on AI models that analyse sensor data (vibration, temperature, consumption) to anticipate faults before they occur and plan the intervention.

Example: a predictive maintenance model warns 10 days in advance of an imminent compressor failure, avoiding an unplanned shutdown of the entire line.

Precision and Recall

Two complementary classification metrics: precision measures what proportion of the model's alerts are correct, and recall measures what proportion of actual cases it manages to detect.

Example: in leak detection, recall is prioritised (not letting any real leak slip through) accepting somewhat lower precision (a few more false alarms).

Perplexity

A metric that evaluates a language model by measuring how surprised it is when predicting a text; the lower the perplexity, the better the model anticipates the next word.

Example: When comparing two LLMs fine-tuned with plant documentation, the team chooses the one with lower perplexity on its technical manuals because it better predicts their vocabulary.

Principal Component AnalysisPCA

A statistical dimensionality-reduction technique that transforms correlated variables into a smaller set of principal components, ordered by the variance they explain, retaining as much information as possible.

Example: With PCA, dozens of correlated sensor signals from a motor are compressed into a few components that summarise its condition and make deviations easier to detect.

Particle Swarm Optimization (PSO)

A population-based optimisation algorithm in which a set of particles moves through the search space, adjusting its trajectory based on its own best position and the swarm's best position.

Example: PSO is applied to tune the controllers of a water treatment plant and minimise its energy consumption.

Prognostics and Health Management (PHM)

A discipline that integrates condition monitoring, fault diagnosis and remaining-life prognosis to decide when to intervene on an asset.

Example: A PHM programme on a refinery's pump fleet prioritises shutdowns according to each unit's estimated deterioration.

Point Cloud

A set of points with 3D coordinates, often with colour or intensity, representing the surface of objects or environments captured with LiDAR, photogrammetry or 3D scanners.

Example: scanning an industrial building with LiDAR to obtain the starting point cloud for the plant's digital twin.

Pose Estimation

Determining an object's position and orientation, or a person's joint positions, from camera images.

Example: A palletising robot calculates each bag's pose on the pallet to grip it at the correct point.

Partial Dependence Plot (PDP)

A curve showing a variable's average effect on a model's prediction, averaging over the rest of the variables. It is used to read an opaque model without opening it.

Example: seeing that predicted specific consumption rises above 8% excess oxygen and using that point as a setpoint limit.

Prescriptive Analytics

Analytics that, besides forecasting what will happen, propose the action to take by combining predictive models with optimisation and the process's real constraints.

Example: the system not only forecasts the electricity demand peak, it also indicates which furnaces to bring forward and which to delay so as not to exceed the contracted power.

Process Mining

A set of techniques that reconstruct, verify and improve a process's real flow from event logs (case identifier, activity and timestamp) extracted from systems such as ERP or MES. It shows how the process really happens, not how it is documented.

Example: Applied to a packaging plant's MES work orders, it reveals that 30 percent of batches pass through quality control twice due to undocumented rework.
Q1 term

Quantization

An optimisation technique that reduces a model's numerical precision (for example from 32 to 8 bits) so it takes up less memory and runs inference faster, with minimal loss of accuracy.

Example: after quantizing the model, a camera with Edge AI runs the inspection on the plant floor itself, without needing a powerful server or sending data to the cloud.
R11 terms

RAGRetrieval-Augmented Generation

A technique that combines an LLM with your own knowledge base. Before answering, the model retrieves relevant information from your documents and uses it to generate the response.

Example: an industrial assistant with RAG can answer questions about your specific plant history, not just general knowledge.

Neural Net

Short form of Neural Network. A computational model made up of layers of interconnected processing units that learn patterns from data.

Example: see the "Neural Network" entry for detail. It is the basic building block of the whole Deep Learning family.

Random Forest

A machine learning algorithm that combines many decision trees trained on random subsets of the data. It tends to give reliable predictions that resist overfitting.

Example: a Random Forest predicts an industrial building's hourly energy consumption using 30 variables (weather, shifts, planned production) with an error under 5%.

RNN (Recurrent Neural Network)

A deep neural network designed for sequential data or time series that retains a "memory" of previous inputs to influence the current output.

Example: In a wind farm, a recurrent network analyses the historical wind-speed sequence to anticipate generation over the next few hours.

Reduced Order Model

A simplified mathematical representation of a complex physical system that reduces computational cost by orders of magnitude while preserving its dominant behaviour.

Example: An energy company integrates a reduced order model of a heat exchanger into its digital twin to simulate it in real time.

RLHFReinforcement Learning from Human Feedback

A training technique in which human evaluators' preferences are used as a reward signal to align a language model's responses with what is considered helpful and safe.

Example: a plant's operator assistant is fine-tuned with RLHF using the technicians' own ratings, so it prioritises safe, actionable responses.

ROC-AUC

A metric that summarises in a single number a classifier's ability to distinguish between classes across all thresholds. It ranges from 0 to 1 and remains reliable even with imbalanced classes.

Example: when comparing two leak-detection models, the team chooses the one with higher ROC-AUC because it better separates leak cases from normal ones without depending on a specific threshold.

Regularization

A set of techniques, such as L1, L2, dropout or early stopping, that penalise a model's complexity during training to prevent overfitting and improve its generalisation to new data.

Example: By adding L2 regularisation, the electricity-consumption prediction model stops memorising one-off historical peaks and performs better in weeks it has never seen.

Reranking

A second retrieval stage in which a more accurate model reorders the initially retrieved documents by relevance, reducing noise before passing them to an LLM. It is a key piece of RAG systems.

Example: After a quick initial search through the plant manuals, a reranker places the exact fault procedure at the top and reduces the assistant's hallucinations.

Remaining Useful Life (RUL)

An estimate of the time (in operating hours, cycles or equivalent units) that an asset can keep running within specification before reaching a failure threshold. It is the central output of asset prognostics models.

Example: Estimating how many cycles a gas turbine bearing has left in order to schedule its replacement at the next planned shutdown.

Recommender System

A information-filtering system that ranks items according to what it predicts will interest a user, based on their previous behaviour or that of similar users.

Example: In an industrial spare-parts catalogue, it suggests to the technician the parts that are usually replaced together with the one just ordered.
S17 terms

SHAPSHapley Additive exPlanations

An explainable AI (XAI) technique that assigns each variable a value indicating how much it contributed to a specific model prediction, making its decisions auditable.

Example: after a quality alarm, SHAP shows that mould temperature and press speed were the two decisive variables, allowing the team to adjust the correct setpoint.

Synthetic Data

Artificial data generated by simulation or AI that mimic real data. They allow models to be trained when real data are scarce, costly or sensitive.

Example: given the lack of examples of a rare fault, synthetic data for that fault are generated so the predictive maintenance model learns to recognise it.

Surrogate Model

A machine learning model that approximates a computationally expensive physical simulation, making it possible to evaluate the system's behaviour at new design points much faster.

Example: A manufacturer trains a surrogate model with CFD results to optimise blade design without repeating thousands of full simulations.

Semantic Segmentation

A computer vision task that classifies each pixel in an image into a predefined category, generating a mask that delimits each class's regions.

Example: In a solar plant, semantic segmentation of drone images labels damaged panels versus operational ones, pixel by pixel.

Self-Supervised Learning

A technique in which the model generates its own labels from unannotated data, learning to predict one part of the input from another. It drastically reduces the need for manual labelling.

Example: a model learns useful representations from millions of unlabelled inspection images, and then just a few annotated examples are enough to specialise it in detecting a specific defect.

Semantic Search

A search technique that interprets the query's intent and meaning, not just the exact words, to return relevant results even if they are phrased differently.

Example: An operator describes motor 3 as vibrating, and semantic search retrieves reports discussing bearing misalignment even though no literal word matches.

Support Vector MachineSVM

A supervised learning algorithm that classifies data by finding the optimal hyperplane that maximises the margin between classes; with kernel functions it also solves non-linear problems.

Example: An SVM classifies transformer oil samples as "normal" or "degraded" from their physicochemical parameters, separating both classes with maximum margin.

Soft Sensor

A software model that estimates in real time a variable that is difficult or costly to measure, based on other easy-to-measure process signals. It is also called an inferential sensor.

Example: Continuously estimating a reactor's product concentration from temperatures, flows and pressures, avoiding the wait for laboratory analysis.

Sensor Fusion

A technique that combines data from several sensors to obtain a more accurate estimate with less uncertainty than any of them would offer individually.

Example: In a wind turbine, accelerometers, temperature sensors and SCADA data are fused to detect an incipient gearbox fault earlier.

State Estimation

A process that infers a system's unmeasured internal variables from noisy or incomplete measurements and a model of the process.

Example: Power grid operators run state estimation every few seconds to reconstruct voltages and flows across the whole network.

Sim-to-Real

An approach that trains models, usually for control or reinforcement learning, in a simulator and transfers them to physical equipment, closing the gap with techniques such as domain randomisation.

Example: training a palletising robot arm's control policy in a digital twin before deploying it on the real robot.

SMOTE (Synthetic Minority Over-sampling Technique)

A technique that balances imbalanced datasets by generating synthetic examples of the minority class through interpolation between nearby neighbours, instead of duplicating existing ones.

Example: generating synthetic fault examples to train a bearing-fault classifier where only 2% of the history are real faults.

Survival Analysis

A family of statistical methods that model the time until an event, such as a failure, handling censored data: assets that have not yet failed by the study's cut-off date.

Example: estimating the probability that a transformer will exceed 5,000 days of service based on the fleet's failure history.

Speech Recognition (ASR)

A technology that converts speech into written text. It is also called ASR or speech-to-text.

Example: An operator dictates the shift report from the plant floor and the system transcribes it directly into the production log.

SCADA (Supervisory Control and Data Acquisition)

A supervisory control and data acquisition system that centralises field signals, displays the process's state and allows action to be taken on it. It is the usual source of the historical data used to train industrial models.

Example: ten years of a wind farm's SCADA history serve as the basis for the model that anticipates gearbox failures.

Spectrogram

A representation of how a signal's energy is distributed across frequencies over time. It converts a time signal into an image that vision networks can classify.

Example: an induced-draft fan's vibrations are converted into a spectrogram and a CNN distinguishes imbalance from misalignment.

Self-Organizing Map (SOM)

An unsupervised neural network, proposed by Kohonen, that projects data with many variables onto a two-dimensional grid while preserving neighbourhood relationships: similar process states end up close together.

Example: placing a glass furnace's operating point on a map where each zone corresponds to a recipe, making the drift towards a rejection region visible.
T7 terms

Transfer Learning

A technique that reuses a model trained in one domain (with lots of data) to solve a similar problem in another domain (with little data). It saves months of training.

Example: starting from a model trained on millions of general images and adapting it to detect specific defects on a production line with only a few hundred of its own images.

Transformer

A neural network architecture that appeared in 2017 and transformed AI, especially in language. It is the foundation of modern LLMs: GPT, Claude, BERT, Llama, Mistral.

Example: the Transformer's attention mechanism allows long sequences to be processed in parallel, making it much faster and more powerful than previous architectures.

Time Series Forecasting

A set of statistical and deep learning techniques (ARIMA, Prophet, LSTM, Temporal Fusion Transformer) used to anticipate future values of variables that evolve over time.

Example: a forecasting model predicts the plant's electricity demand every 15 minutes to optimise energy procurement and reduce the bill.

Tokenization

A preliminary natural language processing step that splits a text into minimal units (tokens), words or fragments, so a model can process it.

Example: before analysing thousands of maintenance reports, the system tokenizes them so the industrial LLM understands the plant's technical terms and abbreviations.

t-SNE (t-Distributed Stochastic Neighbor Embedding)

A non-linear dimensionality-reduction technique that projects high-dimensional data into two or three dimensions while preserving local neighbourhood relationships, useful for visualising clusters.

Example: An engineer visualises thousands of vibration signatures with t-SNE and discovers that misalignment faults form a separate cluster.

Temperature

A parameter that adjusts how sharply or flatly a generative model's probability distribution is shaped when choosing the next token. Low values give more deterministic outputs and high values, more varied ones.

Example: A plant report generator works with a temperature of 0.2 so the figures and format come out the same on every run.

Topic Modeling

A family of unsupervised methods that discover the latent topics of a document collection by grouping words that tend to appear together. LDA is the most widespread algorithm.

Example: grouping 40,000 maintenance reports from a steel plant to see what proportion discuss hydraulic leaks, bearings or electrical shutdowns.
U3 terms

Underfitting

The opposite of overfitting: the model is too simple and fails to capture the data's real patterns. It predicts poorly both in training and in production.

Example: a linear model applied to an industrial process with non-linear relationships will suffer from underfitting. The signal is in the data, but the model doesn't see it.

Uncertainty Quantification

A set of techniques that estimate the confidence level of a model's predictions, delivering a distribution of results instead of a single value. It is key to reliable risk management in critical environments.

Example: A blade's useful-life model accompanies each prediction with its uncertainty quantification, so engineering schedules the inspection with a safety margin based on the estimated confidence.

UMAP (Uniform Manifold Approximation and Projection)

A dimensionality-reduction technique that projects data with many variables into 2 or 3 dimensions while preserving local structure and much of the global structure, faster than t-SNE.

Example: projecting thousands of turbine vibration signatures onto a 2D map where anomalous regimes appear as separate clusters.
V4 terms

Cross-validation

A technique for assessing a model's stability by training it several times with different data partitions, ensuring its performance does not depend on the luck of the split.

Example: instead of splitting the data once into 80/20, it is done 5 times with different splits and averaged. It gives a more reliable measure of expected performance in production.

Vector Database

A database designed to store and index embeddings (numerical vectors) and retrieve them by semantic similarity, not just exact matching. It is the key piece of RAG systems.

Example: an AI assistant instantly retrieves the procedures most similar to an operator's query by searching a vector database with thousands of internal manuals.

Vision Transformer (ViT)

An architecture that applies the Transformer mechanism to images: it divides the image into patches, converts them into a sequence and uses self-attention to capture global relationships. It rivals CNNs in image classification.

Example: a Vision Transformer classifies thermal images of electrical switchboards and detects anomalous hot spots by using the whole image's global context.

Variational Autoencoder (VAE)

A generative autoencoder that learns a probabilistic latent distribution, which makes it possible to generate new data and measure how unusual a sample is.

Example: A VAE trained on a compressor's normal vibrations flags readings with high reconstruction error as anomalous.
W2 terms

Weights

Numerical values that connect a neural network's neurons and are adjusted during training. They are the parameters the model "learns".

Example: an industrial Deep Learning model can have millions or billions of weights. Their combination is what encodes the learned knowledge.

Wavelet Transform

A signal decomposition into components localised in both time and frequency, useful when brief events that the Fourier transform would dilute are of interest.

Example: detecting the impact of a chipped gear tooth within an accelerometer's recording, even if it lasts only milliseconds.
X1 term

XAIExplainable AI

Acronym for Explainable AI. A set of techniques that make a model justify its decisions in an understandable way.

See the full definition underE · Explainable AI.

Y1 term

YOLOYou Only Look Once

A family of real-time computer vision models capable of detecting and classifying multiple objects in an image in a single pass, optimised for very low latency.

Example: a camera with YOLO on the bottling line identifies mispositioned bottles at 60 frames per second and rejects them before labelling.
Z1 term

Zero-shot Learning

A model's ability to solve tasks or recognise classes it never saw during training, relying on general knowledge or auxiliary descriptions.

Example: an LLM classifies maintenance reports into new categories defined only by their name and a brief description, without a single labelled example.

No results

No terms match your search. Try a different keyword.