jablonkagroup/corral-QAs-reports
收藏资源简介:
该数据集是Corral – QA Reports,是Corral集合的一部分,随论文《AI科学家在不科学推理的情况下产生结果》发布。它包含了用于测试模型在Corral环境中事实知识和推理能力的问题-答案评估的模型完成情况和报告。数据集组织为50个配置,每个配置对应可用的环境、模型和评估维度(知识或推理)的组合。例如,一个配置编码了某个模型在给定Corral环境的知识聚焦或推理聚焦QA集上生成的报告。这些完成情况对应于Corral研究中报告的项目反应理论(IRT)分析中使用的QA项目,其中基础知识和推理QA作为潜在知识和推理因素的指标。该资源旨在用于评估、心理测量建模和分析科学代理能力,而非用于通用模型预训练。
This dataset is part of the *Corral* collection accompanying the paper [*AI scientists produce results without reasoning scientifically*](https://arxiv.org/abs/2604.18805). It contains the **model completions and reports** for the **question-answer evaluations** used to test the **factual knowledge** and **reasoning ability** of models across *Corral* environments. The dataset is organized into **50 configurations**, with one configuration for each available combination of **environment**, **model**, and evaluation dimension (**knowledge** or **reasoning**). For example, a config encodes the reports generated by one model on either the knowledge-focused or reasoning-focused QA set for a given Corral environment. These completions correspond to the QA items used in the **Item Response Theory (IRT)** analyses reported in the *Corral* study, where the underlying knowledge and reasoning QAs serve as indicators for the latent **knowledge** and **reasoning** factors. This resource is intended for evaluation, psychometric modeling, and analysis of scientific-agent capabilities rather than for general-purpose model pre-training.




