遇见数据集

Instruction Datasets for Process Mining

收藏
Zenodo2025-05-23 更新2026-05-29 收录
官方服务:

资源简介:

Created based on Process Behavior Corpus and Benchmarking Datasets. Instruction templates were applied to initial datasets, resulting in the following outputs: A-SAD_instructions.csv: An instruction dataset to assess the following task: Given a trace σ, decide if σ is a valid execution of the underlying process or not, without knowing the behavior allowed in the process. T-SAD_instructions.csv: An instruction dataset to assess the following task: Given an eventually-follows relation ef = a ≺ b of a trace σ, decide if ef represents a valid execution order of the two activities a and b that are executed in a process or not, without knowing the behavior allowed in the process. S-NAP_instruction.csv: An instruction dataset to assess the following task: Given an event log L and a prefix p_k of length k, with 1 < k, predict the next activity a_k+1. S-DFD_instructions.csv: An instruction dataset to assess the following task: Given a set of possible activities, generate a difectly follows graph that captures the trace semantics of the process model. S-PTD_instruction.csv: An instruction dataset to assess the following task: Given a set of possible activities, generate a simple process tree that captures the trace semantics of the process model. Each row contains columns: id - process model id that the current sample is derived from; unique_activities - set of possible activities in the current process; instruction - natural-language instruction template filled with the current sample; output - expected answer of the model, according to the given instruction; instruction_type - marks which type of instruction was applied from the perspective of inversion: normal (not inverted), neg_inv (negative inversion, i.e., represents invalid behavior), pos_inv (positive inversion, i.e., represents valid behavior). A-SAD, T-SAD, S-NAP allow positive and negative inversion; S-DFD allows negative inversion; S-PTD does not allow inversion; variant - version of prompt formulation (1-6 for normal instruction_type; 1-4 for neg_inv and pos_inv instruction_type). Additionally, each dataset contains columns for more context: A-SAD: eventually_follows (eventually-follows relation), is_valid (indicates wether the two activities of the relation were executed in a valid order (True) or in an invalid order (False)); T-SAD: trace, is_valid (indicates whether the trace represents a valid execution of the underlying process (True or False)); S-NAP: prefix (ongoing execution of the process), trace (full execution of the process), next (indicates the activity that should be performed next after the last activity of the prefix according to the trace from which the prefix was generated); S-DFD: dfg (directly follows graph that captures the trace semantics of the process model); S-PTD: pt (process tree that captures the trace semantics of the process model). Repository with the related code for applying instruction tuning to LLMs: github.

提供机构:
Zenodo
创建时间:
2025-05-23
二维码
社区交流群
二维码
科研交流群
商业服务