jablonkagroup/questions4manual_annotation
收藏资源简介:
该数据集是Corral集合的一部分,伴随论文《AI科学家在不进行科学推理的情况下产生结果》。它包含从Corral基准测试中选出的用于手动标注认知模式的评估轨迹。数据集按68个配置组织,每个配置对应一个选定的模型、环境、范围和智能体轨迹设置。每个配置中的行是从完整的Corral轨迹集合中分层采样的轨迹。这些轨迹被选中是因为LLM标注器识别出它们为智能体未进行科学推理的案例,因此成为下游人工审查和认知模式分析的目标子集。该资源旨在用于科学智能体行为的标注、定性分析和过程级研究,而非通用模型预训练。
This dataset is part of the Corral collection accompanying the paper AI scientists produce results without reasoning scientifically. It contains the evaluation traces selected for manual annotation of epistemic patterns across the Corral benchmark. The dataset is organized into 68 configurations, with one config per selected model, environment, scope, and agent trace setting. Within each config, rows contain traces sampled in a stratified manner from the full Corral trace collection. The included traces were selected because an LLM annotator identified them as cases where the agents do not reason scientifically, making them a targeted subset for downstream human review and epistemic-pattern analysis. This resource is intended for annotation, qualitative analysis, and process-level study of scientific-agent behaviour rather than for general-purpose model pre-training.




