Additional Tennessee Eastman Process Simulation Data for Anomaly Detection Evaluation
收藏资源简介:
User Agreement, Public Domain Dedication, and Disclaimer of Liability. By accessing or downloading the data or work provided here, you, the User, agree that you have read this agreement in full and agree to its terms. The person who owns, created, or contributed a work to the data or work provided here dedicated the work to the public domain and has waived his or her rights to the work worldwide under copyright law. You can copy, modify, distribute, and perform the work, for any lawful purpose, without asking permission. In no way are the patent or trademark rights of any person affected by this agreement, nor are the rights that any other person may have in the work or in how the work is used, such as publicity or privacy rights. Pacific Science & Engineering Group, Inc., its agents and assigns, make no warranties about the work and disclaim all liability for all uses of the work, to the fullest extent permitted by law. When you use or cite the work, you shall not imply endorsement by Pacific Science & Engineering Group, Inc., its agents or assigns, or by another author or affirmer of the work. This Agreement may be amended, and the use of the data or work shall be governed by the terms of the Agreement at the time that you access or download the data or work from this Website. Description This dataverse contains the data referenced in Rieth et al. (2017). Issues and Advances in Anomaly Detection Evaluation for Joint Human-Automated Systems. To be presented at Applied Human Factors and Ergonomics 2017. Each .RData file is an external representation of an R dataframe that can be read into an R environment with the 'load' function. The variables loaded are named ‘fault_free_training’, ‘fault_free_testing’, ‘faulty_testing’, and ‘faulty_training’, corresponding to the RData files. Each dataframe contains 55 columns: Column 1 ('faultNumber') ranges from 1 to 20 in the “Faulty” datasets and represents the fault type in the TEP. The “FaultFree” datasets only contain fault 0 (i.e. normal operating conditions). Column 2 ('simulationRun') ranges from 1 to 500 and represents a different random number generator state from which a full TEP dataset was generated (Note: the actual seeds used to generate training and testing datasets were non-overlapping). Column 3 ('sample') ranges either from 1 to 500 (“Training” datasets) or 1 to 960 (“Testing” datasets). The TEP variables (columns 4 to 55) were sampled every 3 minutes for a total duration of 25 hours and 48 hours respectively. Note that the faults were introduced 1 and 8 hours into the Faulty Training and Faulty Testing datasets, respectively. Columns 4 to 55 contain the process variables; the column names retain the original variable names. Acknowledgments. This work was sponsored by the Office of Naval Research, Human & Bioengineered Systems (ONR 341), program officer Dr. Jeffrey G. Morrison under contract N00014-15-C-5003. The views expressed are those of the authors and do not reflect the official policy or position of the Office of Naval Research, Department of Defense, or US Government.
用户协议、公共领域贡献声明与免责声明 若您访问或下载本页面提供的数据集或相关作品,即视为您已完整阅读本协议并同意其全部条款。本数据集或作品的所有者、创作者或贡献者已将该作品贡献至公共领域,并依据全球版权法放弃其对该作品的全部权利。您可出于任何合法目的复制、修改、分发或使用该作品,无需获得许可。本协议不影响任何个人的专利权或商标权,亦不影响其他个人对该作品或其使用方式(如肖像权、隐私权)所享有的权利。 太平洋科学与工程集团有限公司(Pacific Science & Engineering Group, Inc.)及其代理方、受让人不对该作品作出任何明示或默示担保,并在法律允许的最大范围内免除因使用该作品而产生的一切责任。您在使用或引用该作品时,不得暗示太平洋科学与工程集团有限公司(Pacific Science & Engineering Group, Inc.)及其代理方、受让人,或该作品的其他作者、确认方对此予以背书。 本协议可随时修订,您访问或下载本网站的数据集或作品时,相关使用行为应受当时有效的协议条款约束。 ### 数据集说明 本数据仓库收录了Rieth等人(2017年)发表的《联合人机自动化系统的异常检测评估:问题与进展》中的相关数据,该成果将在2017年应用人类因素与工效学大会(Applied Human Factors and Ergonomics 2017)上进行展示。 每个.RData文件均为R数据框的外部存储形式,可通过`load`函数加载至R运行环境中。加载后的数据变量分别为`fault_free_training`(无故障训练集)、`fault_free_testing`(无故障测试集)、`faulty_testing`(带故障测试集)与`faulty_training`(带故障训练集),与各.RData文件一一对应。 每个数据框包含55列: 1. 第1列(`faultNumber`,故障编号):在带故障数据集的取值范围为1至20,代表TEP中的故障类型;无故障数据集仅包含故障0,即正常运行工况。 2. 第2列(`simulationRun`,仿真运行编号):取值范围为1至500,代表用于生成完整TEP数据集的不同随机数生成器状态(注:用于生成训练集与测试集的实际随机种子无重叠)。 3. 第3列(`sample`,样本编号):训练数据集的取值范围为1至500,测试数据集的取值范围为1至960。TEP变量(第4至55列)的采样间隔为3分钟,训练集与测试集的总采样时长分别为25小时与48小时。需注意:带故障训练集与带故障测试集分别在运行至1小时与8小时时引入故障。 4. 第4至55列为过程变量,其列名保留原始命名。 ### 致谢 本研究由海军研究办公室(Office of Naval Research)人类与生物工程系统项目(ONR 341)资助,项目合同编号为N00014-15-C-5003,项目主管为Jeffrey G. Morrison博士。本研究表达的观点仅代表作者,不代表海军研究办公室(Office of Naval Research)、美国国防部或美国政府的官方政策或立场。




