K-Forge Evaluation Dataset: When Distilled Knowledge Beats Retrieval for Agents
收藏官方服务:
资源简介:
Evaluation dataset for the K-Forge methodology.72 Bloom-stratified questions across 3 production courses, 7,818 model answers from 10 LLMs under 3 knowledge-representation configs (K-Forge structured BoK, standard RAG, ablation), and 28,593 judgments from a cross-vendor tri-judge panel (GPT-4o, Claude Sonnet, Qwen). Includes verbatim QA scores, reference-free mastery pairwise comparisons, budget-parity controls, and per-question refusal analysis. Raw course content excluded (copyright); context character counts provided for reproducibility.
提供机构:
Zenodo创建时间:
2026-06-10



