jablonkagroup/corral-QAs-topic_reports
收藏资源简介:
该数据集是Corral集合的一部分,伴随论文《AI科学家在不进行科学推理的情况下产生结果》。它包含了用于测试模型在所有8个Corral环境中的事实知识和推理能力的问答评估的平均结果。数据集分为48个配置,每个配置对应环境、模型和评估维度(知识或推理)的组合。在每個配置中,行汇总了相应知识或推理评估的平均问答结果。这些汇总的问答报告总结了Corral研究中项目反应理论(IRT)分析中使用的项目结果,支持潜在的知识和推理因素。该资源旨在用于评估、心理测量建模和科学代理能力的比较分析,而非通用模型预训练。
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the averaged results of the question-answer evaluations used to test the factual knowledge and reasoning ability of models across all 8 Corral environments. The dataset is organized into 48 configurations, with one config per environment, model, and evaluation dimension combination. Within each config, rows summarize averaged QA outcomes for the corresponding knowledge or reasoning evaluation. These aggregated QA reports summarize the outcomes of the items used in the Item Response Theory (IRT) analyses reported in the Corral study, where they support the latent knowledge and reasoning factors. This resource is intended for evaluation, psychometric modeling, and comparative analysis of scientific-agent capabilities rather than for general-purpose model pre-training.




