遇见数据集

Batch Distillation Data for Developing Machine Learning Anomaly Detection Methods

收藏
Zenodo2026-05-28 更新2026-05-26 收录
官方服务:

资源简介:

This database provides a resource for machine-learning-based anomaly detection (AD) in chemical processes. It includes data from 119 experimental runs on a laboratory-scale batch distillation plant across diverse operating conditions and mixtures, with paired fault-free runs and deliberately induced anomalies. For 115 of these experiments, corresponding simulation predictions are provided, creating a hybrid experimental-simulation dataset. The data include multivariate time-series from sensors and actuators with measurement-uncertainty estimates, complemented by unconventional modalities such as online benchtop NMR concentration profiles, image and audio recordings. Every anomalous experiment is richly documented with extensive metadata and expert annotations capturing both the presence and causes of anomalies. The simulation results are released together with detailed simulation metadata; they cover fault-free operation and anomalies that can be represented within the simulation model, particularly anomalies induced by setpoint perturbations. The data are organized in a structured, ready-to-use format and made freely available to support benchmarking and the development of advanced AD methods. By linking anomalies to their underlying causes and corresponding simulation results, the database also enables research on interpretable and explainable machine learning (ML), as well as strategies for anomaly mitigation.

本数据库为化工过程中基于机器学习的异常检测(Anomaly Detection, AD)任务提供了研究资源。该数据集涵盖实验室规模间歇精馏装置在多样化操作工况与物料体系下的119组实验运行数据,包含正常无故障运行组与人为引入异常的配对实验样本。针对每组实验,本数据集提供了来自传感器与执行器的多变量时间序列数据(附带测量不确定性估计值),同时补充了非常规模态数据:在线台式核磁共振(Nuclear Magnetic Resonance, NMR)浓度谱、图像与音频录制文件。每一组异常实验均配有详尽的元数据与专家标注信息,该元数据可同时记录异常的发生情况与成因。数据集采用结构化、可直接使用的格式进行组织,并免费开放以供先进异常检测方法的基准测试与算法开发使用。通过将异常事件与其潜在成因相关联,本数据库还可支撑可解释与可阐释机器学习(Machine Learning, ML)以及异常缓解策略相关研究。

提供机构:
Zenodo
创建时间:
2026-04-12
二维码
社区交流群
二维码
科研交流群
商业服务