Public benchmark dataset for Conformance Checking in Process Mining
收藏资源简介:
This dataset contains a variety of publicly available real-life event logs. We derived two types of Petri nets for each event log with two state-of-the-art process miners : Inductive Miner (IM) and Split Miner (SM). Each event log-Petri net pair is intended for evaluating the scalability of existing conformance checking techniques.We used this data-set to evaluate the scalability of the S-Component approach for measuring fitness. The dataset contains tables of descriptive statistics of both process models and event logs. In addition, this dataset includes the results in terms of time performance measured in milliseconds for several approaches for both multi-threaded and single-threaded executions. Last, the dataset contains a cost-comparison of different approaches and reports on the degree of over-approximation of the S-Components approach. The description of the compared conformance checking techniques can be found here: https://arxiv.org/abs/1910.09767. <br><br>Update:The dataset has been extended with the event logs of the BPIC18 and BPIC19 logs. BPIC19 is actually a collection of four different processes and thus was split into four event logs. For each of the additional five event logs, again, two process models have been mined with inductive and split miner. We used the extended dataset to test the scalability of our tandem repeats approach for measuring fitness. The dataset now contains updated tables of log and model statistics as well as tables of the conducted experiments measuring execution time and raw fitness cost of various fitness approaches. The description of the compared conformance checking techniques can be found here: https://arxiv.org/abs/2004.01781.<br>Update: <br>The dataset has also been used to measure the scalability of a new Generalization measure based on concurrent and repetitive patterns. : A concurrency oracle is used in tandem with partial orders to identify concurrent patterns in the log that are tested against parallel blocks in the process model. Tandem repeats are used with various trace reduction and extensions to define repetitive patterns in the log that are tested against loops in the process model. Each pattern is assigned a partial fulfillment. The generalization is then the average of pattern fulfillments weighted by the trace counts for which the patterns have been observed. The dataset no includes the time results and a breakdown of Generalization values for the dataset.<br> <br>
本数据集包含多种公开可用的真实世界事件日志(event log)。我们采用两款当前前沿的流程挖掘工具——归纳挖掘器(Inductive Miner,IM)与拆分挖掘器(Split Miner,SM)——为每份事件日志生成两类Petri网。每一组事件日志-Petri网配对均用于评估现有一致性检查(conformance checking)技术的可扩展性。我们曾使用该数据集评估用于衡量适配度的S-Component方法的可扩展性。 本数据集收录了流程模型与事件日志的描述性统计表。此外,数据集还包含多线程与单线程执行场景下,多种方法的毫秒级时间性能测试结果。最后,数据集还涵盖了不同方法的成本对比报告,以及S-Component方法的过近似程度相关分析数据。所对比的一致性检查技术的详细说明可参见:https://arxiv.org/abs/1910.09767。 更新:本数据集已扩展加入BPIC18与BPIC19事件日志。其中BPIC19实际包含四个独立流程,因此被拆分为四份事件日志。针对新增的五份事件日志,我们同样使用归纳挖掘器与拆分挖掘器各生成了两份流程模型。我们依托扩展后的数据集测试了用于衡量适配度的串联重复(tandem repeats)方法的可扩展性。当前数据集已更新包含日志与模型统计信息表,以及针对多种适配度方法的执行时间与原始适配成本的实验测试表。所对比的一致性检查技术的详细说明可参见:https://arxiv.org/abs/2004.01781。 更新:本数据集还被用于评估一种基于并发与重复模式的新型泛化度量方法的可扩展性:该方法将并发预言机(concurrency oracle)与偏序集结合,用于识别日志中的并发模式,并与流程模型中的并行块进行比对;同时将串联重复序列与多种轨迹约简及扩展方法结合,用于定义日志中的重复模式,并与流程模型中的循环结构进行比对。每个模式会被赋予部分满足度,最终泛化度量值为各模式满足度按观测到该模式的轨迹数量加权后的平均值。本数据集现已补充收录了该泛化度量的时间结果与细分数据。




