jablonkagroup/corral-QAs
收藏资源简介:
该数据集是Corral集合的一部分,包含用于测试模型在8个Corral环境中事实知识和推理能力的问题-答案对。数据集按16个配置组织,对应8个环境和两个评估维度(知识和推理)的笛卡尔积。这些QA对用于项目反应理论分析,作为潜在知识和推理因素的指标。数据集主要用于评估、心理测量建模和科学智能体能力分析,而非通用模型预训练。
This dataset is part of the Corral collection, containing question-answer pairs for testing models' factual knowledge and reasoning capabilities across 8 Corral environments. It is organized into 16 configurations, which correspond to the Cartesian product of the 8 environments and two evaluation dimensions: knowledge and reasoning. These QA pairs are used for Item Response Theory (IRT) analysis, serving as indicators of latent knowledge and reasoning factors. The dataset is primarily intended for evaluation, psychometric modeling, and scientific AI Agent capability analysis, rather than general-purpose model pre-training.




