遇见数据集

Multimodal Process Monitoring with On-Demand Disambiguation

收藏
Zenodo2026-05-22 更新2026-05-26 收录
官方服务:

资源简介:

This online appendix includes the supplementary material referenced in the article "Multimodal Process Monitoring with On-Demand Disambiguation". In addition to the evaluation results of the proposed MMPMOD architecture and vision model, the material provides the data and scripts required to inspect the experimental evaluation and reproduce the ablation study results reported in the article. The provided files allow inspection of the raw disambiguation attempts, including captured images, disambiguation results, and ground truth, enabling independent verification of the reported evaluation results. Specifically, the material includes the following artifacts: 1. The results of the experimental evaluation described in the article, including the evaluation of the vision model and the end-to-end evaluation results with the ablation study: The vision model has been trained and tested using the publicly available dataset available here. The results of the vision model evaluation are included in file Vision_model_evaluation_results.zip. The end-to-end evaluation is based on 50 enactments of a phlebotomy process in an IoT-monitored lab setup, in which an implementation of the MMPMOD architecture is integrated to disambiguate events during process monitoring by complementing IoT sensing with a video sensing modality. For the end-to-end evaluation, the implementation of the MMPMOD architecture has been deployed on a MacBook Pro M3 Max with 36GB RAM, connected with a Logitech StreamCam camera with a 1920x1080 resolution. Disambiguation in the end-to-end evaluation was based on an aggregation window of 10 camera images. The results of the end-to-end evaluation are included in file Disambiguation_results.xlsx. In the file, the results are organized as follows: Sheet "End-to-end evaluation results" presents, for each disambiguation attempt, whether it is a genuine disambiguation attempt or a false ambiguity, the actual enactment, the disambiguation result, whether the result matches the actual enactment, and the time required for the inference by the vision model. In this sheet, each row corresponds to one disambiguation attempt. Sheet "End-to-end evaluation notes" presents manual notes taken by an annotator during process enactment, indicating genuine and false ambiguities, misclassifications, and misclassification types. In this sheet, each row corresponds to one process enactment. Sheet "Ablation study results" presents the results of an ablation study, in which the implemented disambiguation approach is compared against a null model approach, a vision-based approach without temporal aggregation, and multiple vision-model approaches with varying temporal aggregation. The null model approach does not rely on the vision sensing modality and lacks contextualization; therefore, it disambiguates events by randomly selecting one of the possible interpretations. The vision-based approach without temporal aggregation relies on a single camera image to perform the disambiguation attempt. The remaining approaches vary the number of camera images used for temporal aggregation, ranging from 2 to 5 images. For each approach, the sheet reports correctness of each disambiguation attempt, aggregated disambiguation accuracy, and, except for the null model approach, inference time (excluding camera activation and image acquisition). In this sheet, each row corresponds to one disambiguation attempt. 2. Three event logs, which have been encoded using the XES standard to facilitate inspection with common process mining tools. The three logs, which are included in file Logs.zip, are: Ambiguous_log_standard.xes: contains traces with ambiguous events derived from IoT event abstraction in the phlebotomy monitoring setting. Disambiguated_log_standard.xes: contains traces derived from the disambiguation attempts by the MMPMOD implementation. Ground_truth_log_standard.xes: contains traces describing the actual executions of the monitored process, derived manually by a human annotator observing the process enactments. 3. All the images captured by the MMPMOD implementation when attempting disambiguation during the 50 process enactments of the end-to-end evaluation. The images are included in file Captured_images.zip. In the compressed file, the 10 images captured for each disambiguation attempt are located in an individual subfolder. Because during the end-to-end evaluation disambiguation was attempted 259 times, the compressed file contains 259 subfolders. 4. In the spirit of reproducibility, file ablation_study.py encodes a Python script that allows the ablation study to be executed again using the images in Captured_images.zip. To execute the script, place it in the directory java/ch/unisg/mmpmod/ambiguityresolver/ml of the MMPMOD implementation, which can be found here. The script reuses the vision inference component and model included in the implementation. Moreover, the uncompressed Captured_images.zip must be included in the same location.

提供机构:
Zenodo
创建时间:
2026-03-17
二维码
社区交流群
二维码
科研交流群
商业服务