2077AIDataFoundation/KINA
收藏资源简介:
KINA(诺亚方舟知识索引)是一个多学科知识基准测试数据集,旨在评估大语言模型是否能够解决高密度、有来源依据的研究生级别问题,覆盖广泛的学科领域。该数据集包含899个十选项的伪多项选择题,涵盖261个细粒度子领域、70个领域和12个顶级学科,包括工程学、科学、医学、法学、文学与艺术、农学、经济学、教育学、管理学、社会学、历史和哲学。每个问题均用英文编写,提供10个答案选项(A到J),并包含正确答案、选项级解释、来源材料和学科元数据。数据集强调学科代表性、高质量控制和有界测试预算下的排名稳定性,通过四阶段质量控制流程确保问题质量,适用于模型评估和研究。
KINA (Knowledge Index of Noahs Ark) is a multidisciplinary knowledge benchmark for evaluating whether large language models can solve high-density, source-grounded, graduate-level questions across a broad map of human disciplines. The dataset contains 899 ten-option pseudo-multiple-choice questions covering 261 fine-grained subfields, 70 fields, and 12 top-level disciplines, including Engineering, Science, Medicine, Law, Literature and Arts, Agronomy, Economics, Education, Management, Sociology, History, and Philosophy. Each question is written in English, has 10 answer options (A-J), and includes the correct answer, option-level explanations, source materials, and discipline metadata. The benchmark emphasizes disciplinary representativeness, quality control under reviewer incentives, and ranking stability under bounded test budgets, with a four-stage quality-control pipeline to ensure question quality, suitable for model evaluation and research.




