Batch Distillation Data for Developing Machine Learning Anomaly Detection Methods
收藏资源简介:
This database provides a resource for machine-learning-based anomaly detection (AD) in chemical processes. It includes data from 119 experimental runs on a laboratory-scale batch distillation plant across diverse operating conditions and mixtures, with paired fault-free runs and deliberately induced anomalies. For 115 of these experiments, corresponding simulation predictions are provided, creating a hybrid experimental-simulation dataset. The data include multivariate time-series from sensors and actuators with measurement-uncertainty estimates, complemented by unconventional modalities such as online benchtop NMR concentration profiles, image and audio recordings. Every anomalous experiment is richly documented with extensive metadata and expert annotations capturing both the presence and causes of anomalies. The simulation results are released together with detailed simulation metadata; they cover fault-free operation and anomalies that can be represented within the simulation model, particularly anomalies induced by setpoint perturbations. The data are organized in a structured, ready-to-use format and made freely available to support benchmarking and the development of advanced AD methods. By linking anomalies to their underlying causes and corresponding simulation results, the database also enables research on interpretable and explainable machine learning (ML), as well as strategies for anomaly mitigation.



