遇见数据集

Real and Synthetic Data for Industrial Anomaly Detection in Injection Molding

收藏
Zenodo2026-04-09 更新2026-05-26 收录
官方服务:

资源简介:

Overview This collection contains a blend of real-world and synthetic datasets designed for industrial anomaly detection. The centerpiece is Real Molding, a dataset collected directly from a real industrial injection molding machine. This real dataset is supported by five synthetic datasets (Lines A through E). These synthetic lines simulate varying operational conditions, from highly stable environments to turbulent regimes, providing a diverse historical pool for transfer learning, frugal AI, and zero-shot anomaly detection research. Data Structure Real Molding Format The real-world dataset captures specific process variables from the injection molding cycle: timestamp: Date and time of measurement. Injection Time: Physical process variable. Plastification Time: Physical process variable. Cycle Time: Physical process variable. Cushion: Physical process variable. Max Pressure: Physical process variable. label: Binary indicator (0 = normal operation, 1 = genuine process deviation anomaly). Synthetic Lines Format The synthetic datasets share a generalized sensor feature space: timestamp: Date and time of measurement. Temperature: Process temperature. Pressure: Process pressure. Elapsed_time (Lines A and B only): Machine runtime. label: Binary indicator (0 = normal operation, 1 = anomaly). Dataset Descriptions Real Molding (Industrial Target): Source: Real industrial injection molding machine. Records: 2,999 production cycles. Features: 5 physical variables. Characteristics: Realistic class imbalance representing genuine process deviations with complex, entangled feature distributions. Anomalies: 92 anomalies (3.06% rate). Size: 83.37KB Line A (Stable/Large): Records: 10,000. Features: 3 variables (Temperature, Pressure, Elapsed Time). Characteristics: Simulates a stable production line with low noise and distinct anomaly peaks, serving as a clean knowledge source. Anomalies: 18 anomalies (0.18% rate). Size: 775.68KB Line B (Balanced): Records: 5,000. Features: 3 variables (Temperature, Pressure, Elapsed Time). Characteristics: Represents a standard baseline with moderate noise levels. Anomalies: 50 anomalies (1.00% rate). Size: 390.71KB Line C (Turbulent): Records: 5,000. Features: 2 variables (Temperature, Pressure). Characteristics: Highly volatile process with significant noise and extreme class overlap. Anomalies: 200 anomalies (4.00% rate). Size: 295.02KB Line D (Noisy/Sparse): Records: 5,000. Features: 2 variables (Temperature, Pressure). Characteristics: Noisy conditions with a very low frequency of anomalies, challenging the detection process. Anomalies: 15 anomalies (0.30% rate). Size: 295.10KB Line E (Clean): Records: 5,000. Features: 2 variables (Temperature, Pressure). Characteristics: Highly controlled process with low noise and well-defined anomaly signatures. Anomalies: 25 anomalies (0.50% rate). Size: 295.08KB Data Statistics Synthetic Datasets Temperature Range Pressure Range Elapsed Time Range % of Anomalies LineA_Stable_10K ~179-180 ~159-160 ~34-35 0.18% LineB_Flux ~188-191 ~19-20 ~19-20 1.00% LineC_Turbulent ~196-210 ~97-103 N/A 4.00% LineD_SpikeControl ~196-202 ~97-102 N/A 0.30% LineE_SmoothRun ~199-200 ~99-100 N/A 0.50% Suggested Applications Zero-shot anomaly detection and transfer learning across heterogeneous feature spaces; Model retrieval and historical model reuse for resource-constrained edge environments (Frugal AI); Time series analysis and handling of severe class imbalances in streaming data; Benchmarking meta-learning and algorithm selection techniques for industrial IoT; Comparative analysis of domain shifts between synthetic proxies and real-world industrial targets. Contact Davide Carneiro davide.r.carneiro@inesctec.pt Escola Superior de Tecnologia e Gestão, Instituto Politécnico do Porto, 4610-156 Felgueiras, Portugal INESC TEC, R. Dr. Roberto Frias, 4200-465 Porto, Portugal

数据集概述 本数据集集合融合了真实世界与合成数据集,专为工业异常检测任务设计。其核心为**真实注塑(Real Molding)**数据集,该数据集直接采集自真实工业注塑机。该真实数据集配套有5个合成数据集(A至E行)。这些合成数据集模拟了从高度稳定到剧烈波动的各类运行工况,为迁移学习、节俭AI(Frugal AI)以及零样本(Zero-shot)异常检测研究提供了多样化的历史数据集池。 ## 数据结构 ### 真实注塑数据集格式 该真实数据集采集了注塑循环中的特定过程变量: - 时间戳:测量的日期与时间 - 注塑时间:物理过程变量 - 塑化时间:物理过程变量 - 循环周期:物理过程变量 - 料垫:物理过程变量 - 最大压力:物理过程变量 - 标签:二进制标记(0 = 正常运行,1 = 真实工艺偏差异常) ### 合成数据集格式 所有合成数据集共享通用的传感器特征空间: - 时间戳:测量的日期与时间 - 温度:工艺温度 - 压力:工艺压力 - 运行时长(仅A、B两行包含):机器运行时长 - 标签:二进制标记(0 = 正常运行,1 = 异常) ## 数据集详情 ### 真实注塑数据集(工业目标数据集) - 来源:真实工业注塑机 - 样本量:2999个生产循环 - 特征数:5个物理变量 - 特性:具备真实的类别不平衡性,特征分布复杂且相互纠缠,还原真实工艺偏差场景 - 异常样本数:92个,占比3.06% - 大小:83.37KB ### 数据集A(稳定型大样本) - 样本量:10000 - 特征数:3个变量(温度、压力、运行时长) - 特性:模拟低噪声、异常峰值显著的稳定生产线,可作为干净的知识源 - 异常样本数:18个,占比0.18% - 大小:775.68KB ### 数据集B(平衡型) - 样本量:5000 - 特征数:3个变量(温度、压力、运行时长) - 特性:代表中等噪声水平的标准基线场景 - 异常样本数:50个,占比1.00% - 大小:390.71KB ### 数据集C(湍流型) - 样本量:5000 - 特征数:2个变量(温度、压力) - 特性:高波动工艺场景,伴随显著噪声与严重的类别重叠 - 异常样本数:200个,占比4.00% - 大小:295.02KB ### 数据集D(噪声稀疏型) - 样本量:5000 - 特征数:2个变量(温度、压力) - 特性:高噪声环境且异常频率极低,对异常检测算法构成挑战 - 异常样本数:15个,占比0.30% - 大小:295.10KB ### 数据集E(纯净型) - 样本量:5000 - 特征数:2个变量(温度、压力) - 特性:高度可控的低噪声工艺,异常特征定义清晰 - 异常样本数:25个,占比0.50% - 大小:295.08KB ## 数据统计信息 | 合成数据集 | 温度范围 | 压力范围 | 运行时长范围 | 异常占比 | |------------------|----------|----------|--------------|----------| | LineA_Stable_10K | ~179-180 | ~159-160 | ~34-35 | 0.18% | | LineB_Flux | ~188-191 | ~19-20 | ~19-20 | 1.00% | | LineC_Turbulent | ~196-210 | ~97-103 | 无此项 | 4.00% | | LineD_SpikeControl| ~196-202 | ~97-102 | 无此项 | 0.30% | | LineE_SmoothRun | ~199-200 | ~99-100 | 无此项 | 0.50% | ## 推荐应用场景 1. 异构特征空间下的零样本异常检测与迁移学习; 2. 资源受限边缘环境下的模型检索与历史模型复用(节俭AI); 3. 时序数据分析与流式数据中严重类别不平衡问题的处理; 4. 面向工业物联网的元学习与算法选择技术的基准测试; 5. 合成代理数据集与真实工业目标数据集间的域偏移对比分析。 ## 联系方式 Davide Carneiro 邮箱:davide.r.carneiro@inesctec.pt 葡萄牙波尔图理工学院高等技术与管理学院,4610-156 费尔盖拉斯 INESC TEC研究院,罗伯托·弗里亚斯博士路4200-465 波尔图,葡萄牙

提供机构:
Zenodo
创建时间:
2026-04-09
二维码
社区交流群
二维码
科研交流群
商业服务