Phase 2 Synthetic Event Logs for Counterfactual Process Dynamics in BPIC-AF v75
收藏资源简介:
This dataset contains the synthetic event logs generated during Phase 2 (stochastic generative phase) of the BPIC-AF v75 experimental pipeline. The logs were produced under a controlled simulation protocol designed to approximate counterfactual process dynamics while preserving structural and statistical constraints inferred during Phase 1 prior estimation. The generation mechanism implements a probabilistic process model defined over: Empirical activity transition kernels Resource-conditional transition matrices Case-level duration distributions Inter-arrival stochastic processes Boundary forcing derived from observed arrival intensity Each synthetic trace is constructed via sequential sampling from a constrained state-transition model: P(at+1,rt+1,Δtt+1∣at,rt,θ)P(a_{t+1}, r_{t+1}, \Delta t_{t+1} \mid a_t, r_t, \theta)P(at+1,rt+1,Δtt+1∣at,rt,θ) where: ata_tat denotes the activity at step ttt, rtr_trt denotes the executing resource, Δtt\Delta t_tΔtt denotes inter-event temporal increments, θ\thetaθ represents the parameter vector estimated in Phase 1. The generative process enforces: Empirical marginal preservation of activity frequencies. Controlled perturbation of transition entropy. Stochastic consistency of resource allocation dynamics. Temporal coherence via sampled duration models. Reproducibility through deterministic seed control. The dataset includes, per scenario and replication: CSV-formatted event logs (t1.csv, t2.csv) Optional XES exports (process mining compliant) Metadata descriptors (meta.json) Run-level manifest for audit traceability (phase2_manifest.csv) Each trace contains: case_id activity timestamp resource Additional synthetic or preserved attributes depending on configuration. The dataset is suitable for: Process mining benchmarking Transition structure robustness analysis Counterfactual intervention simulation Statistical physics–based network topology evaluation Digital twin validation experiments All logs are generated without inclusion of personally identifiable information and are fully synthetic, derived from probabilistic abstractions of the original empirical process. The generation protocol ensures structural separation between: The observational model (sensor layer) The generative digital twin The evaluation and control modules This separation guarantees methodological non-circularity and supports reproducible Q1-level experimental validation workflows.



