Prabhjotschugh/firstpass-peer-review
收藏资源简介:
FirstPass是一个多领域、多轮次的科学同行评审数据集,基于《自然通讯》的透明同行评审文件构建。它包含来自五个科学领域的3668个完整的同行评审对话,并带有从实际修订周期结果中推导出的编辑结果标签。该数据集涵盖生物学、化学、神经科学、物理学和地球科学领域,数据来源于《自然通讯》期刊(ISSN 2041-1723),时间范围为2023年1月至2025年12月。数据集包含三种配置:用于结果预测的分类分割(cls)、用于生成任务(如评论生成和评审员更新)的监督微调分割(sft),以及包含完整解析记录(如论文内容、所有评审轮次和元数据)的全结构化记录(records)。数据经过严格的质量过滤和完整性审核,适用于文本分类、文本生成和问答等NLP任务,特别是同行评审相关的应用。
FirstPass is a multi-domain, multi-turn scientific peer review dataset constructed based on the transparent peer review documents of *Nature Communications*. It contains 3668 complete peer review conversations across five scientific disciplines, paired with editorial outcome labels derived from the results of actual review cycles. This dataset covers the fields of biology, chemistry, neuroscience, physics, and earth science, with data sourced from the journal *Nature Communications* (ISSN 2041-1723), spanning from January 2023 to December 2025. The dataset includes three configurations: the classification split (cls) for outcome prediction, the supervised fine-tuning split (sft) for generation tasks such as comment generation and reviewer updating, and the fully structured record set (records) that contains complete parsed records such as paper content, all review rounds, and metadata. The data has undergone rigorous quality filtering and integrity audits, making it suitable for NLP tasks including text classification, text generation, and question answering, especially for peer review-related applications.




