Benchmark Dataset for Conversational Analysis of OCEL 2.0 Object-Centric Event Logs
收藏资源简介:
This dataset is a benchmark for evaluating conversational interfaces over object-centric event logs following the OCEL 2.0 format. The dataset is derived from a simulated Procure-to-Pay (P2P) object-centric event log based on real SAP transaction semantics and released as part of the OCEL 2.0 reference logs. From this log, we generated a collection of 2,552 natural language question–answer pairs to assess the ability of Large Language Models (LLMs) and Retrieval-Augmented Generation (RAG) pipelines to analyze object-centric process data. The questions cover four complementary perspectives of object-centric process analysis: global process statistics; event-related information; object lifecycle information; and timestamp-based information. Each question is associated with a boolean ground-truth answer (True/False), enabling reproducible quantitative evaluation using standard binary classification metrics. For each positive question, a corresponding negative question was generated by minimally perturbing factual details, resulting in a balanced dataset. The dataset is provided in multiple variants, including: the shuffled complete dataset; and separate files for each analytical perspective (global statistics, events, objects, timestamps). The dataset accompanies the paper “Enabling Natural Language Analysis for Object-Centric Event Logs” by Angelo Casciani, Mario Luca Bernardi, Marta Cimitile, and Andrea Marrella, published in the "Artificial Intelligence and Processes" collection of the Process Science journal.



