KG-Align-RL: Evidence Artifacts (per-checkpoint eval results, classified trajectories, SFT corpus)
收藏资源简介:
Reproducibility evidence artifacts for the KG-Align-RL analysis paper (verifiable process supervision via knowledge graphs for agentic RL; Qwen2.5-7B/14B + GRPO on ComplexWebQuestions/Freebase). These artifacts let a reader re-derive every number in the paper's tables/figures and inspect the qualitative phenomena (step-level peak-then-collapse, the 4-mode failure taxonomy, correct-via-tool counts). Contents: per-checkpoint full-test eval JSONs; 7-category classified trajectory dumps and per-sample classifications; pass@k / self-consistency results; mode-4 reward-decomposition and token-entropy series; the rule-based (no-LLM-teacher) SFT trajectory corpus; temporal/volatility and Category-B analysis JSONs. Not included (third-party / licensed / regenerable): the Freebase KG subgraph (from RoG / Reasoning-on-Graphs) and the verl CWQ parquets — regenerate from source per the repository's DATA.md. Code: see the accompanying (anonymized) repository linked in the paper.



