遇见数据集

Dataset with condition monitoring vibration data annotated with technical language, from paper machine industries in northern Sweden

收藏
data.europa2023-12-22 更新2025-05-24 收录
官方服务:

资源简介:

Labelled industry datasets are one of the most valuable assets in prognostics and health management (PHM) research. However, creating labelled industry datasets is both difficult and expensive, making publicly available industry datasets rare at best, in particular labelled datasets. Recent studies have showcased that industry annotations can be used to train artificial intelligence models directly on industry data ( https://doi.org/10.36001/ijphm.2022.v13i2.3137 , https://doi.org/10.36001/phmconf.2023.v15i1.3507 ), but while many industry datasets also contain text descriptions or logbooks in the form of annotations and maintenance work orders, few, if any, are publicly available. Therefore, we release a dataset consisting with annotated signal data from two large (80mx10mx10m) paper machines, from a Kraftliner production company in northern Sweden. The data consists of 21 090 pairs of signals and annotations from one year of production. The annotations are written in Swedish, by on-site Swedish experts, and the signals consist primarily of accelerometer vibration measurements from the two machines. The dataset is structured as a Pandas dataframe and serialized as a pickle (.pkl) file and a JSON (.json) file. The first column (‘id’) is the ID of the samples; the second column (‘Spectra’) are the fast Fourier transform and envelope-transformed vibration signals; the third column (‘Notes’) are the associated annotations, mapped so that each annotation is associated with all signals from ten days before the annotation date, up to the annotation date; and finally the fourth column (‘Embeddings’) are pre-computed embeddings using Swedish SentenceBERT. Each row corresponds to a vibration measurement sample, though there is no distinction in this data between which sensor or machine part each measurement is from.

带标注的工业数据集是故障预测与健康管理(Prognostics and Health Management,PHM)研究中最具价值的资产之一。然而,构建带标注的工业数据集既困难又成本高昂,导致公开可用的工业数据集本就稀缺,带标注的数据集更是凤毛麟角。 近期研究表明,工业标注可直接用于在工业数据上训练人工智能模型(相关研究链接:https://doi.org/10.36001/ijphm.2022.v13i2.3137、https://doi.org/10.36001/phmconf.2023.v15i1.3507),但尽管多数工业数据集都包含以标注和维修工单形式存在的文本描述或日志记录,公开可用的此类数据集却寥寥无几。 为此,我们发布了一套数据集,其源自瑞典北部一家牛皮纸浆生产企业的两台大型(80米×10米×10米)造纸机的标注信号数据。该数据集包含了为期一年生产过程中产生的21090组信号与标注对。标注由瑞典现场专家以瑞典语撰写,信号则主要来自两台设备的加速度振动测量数据。 本数据集采用Pandas数据框格式存储,并序列化为Pickle(.pkl)与JSON(.json)两种文件格式。各列含义如下:第一列"id"为样本ID;第二列"Spectra"为经快速傅里叶变换与包络变换处理后的振动信号;第三列"Notes"为对应的标注,经映射后,每条标注均关联标注日期前10天至标注当日的所有信号;第四列"Embeddings"为使用瑞典语SentenceBERT预计算得到的词嵌入向量。每一行对应一条振动测量样本,但本数据未区分每条测量数据对应的传感器或机器部件。

创建时间:
2023-12-21
二维码
社区交流群
二维码
科研交流群
商业服务