遇见数据集

K-Forge Evaluation Dataset: When Distilled Knowledge Beats Retrieval for Agents

收藏
Zenodo2026-06-20 更新2026-06-12 收录
官方服务:

资源简介:

Evaluation dataset for the K-Forge methodology.72 Bloom-stratified questions across 3 production courses, 7,818 model answers from 10 LLMs under 3 knowledge-representation configs (K-Forge structured BoK, standard RAG, ablation), and 28,593 judgments from a cross-vendor tri-judge panel (GPT-4o, Claude Sonnet, Qwen). Includes verbatim QA scores, reference-free mastery pairwise comparisons, budget-parity controls, and per-question refusal analysis. Raw course content excluded (copyright); context character counts provided for reproducibility.

提供机构:
Zenodo
创建时间:
2026-06-10
二维码
社区交流群
二维码
科研交流群
商业服务