Flexible Policy Learning for Joint Activity Routing and Resource Allocation in Business Process Simulation
收藏资源简介:
Event logs: AC_CRE (academic credential management) BPIC12_W (loan application subprocess, BPIC 2012) BPIC17_W (loan application subprocess, BPIC 2017). Evaluation policies (K=10 runs each): RA-RR random activity, random resource (lower bound) DM-RR empirical Markov routing, random resource (as-is baseline) DM-GR empirical Markov routing, greedy min-processing-time resource DM-DRL empirical Markov routing, PPO-trained resource-only agent DRL-DRL full joint PPO agent selecting both activity and resource Operational metrics (Evaluated at T95, T90, T75, T50): CR fraction of cases completing within SLA threshold T CIR relative compliance improvement over the original log Similarity metrics (Camargo et al. framework, lower is better): NGD · AED · RED · CED · CWD · CAR · CTD Folders (2×2 Flexibility mask × KL regularization grid): flex/ flexible nucleus mask (top-k/p), no KL regularization flex_kl_reg/ flexible nucleus mask + KL regularization noflex_matrix/ unrestricted mask (top_k=100, top_p=1), no KL regularization no_flex_kl_reg/ unrestricted mask + KL regularization Each folder contains: aggregated (mean ± 95% CI per log × policy), runs_wide (one row per run), runs_long (metric-melted long format). opra_results.csv — all four experimental conditions combined into a single file for cross-condition analysis.



