WENJINLIU/Paper-Review-Dataset
收藏资源简介:
该数据集包含国际学习表征会议(ICLR)2023年、2024年和2025年的论文提交和评审数据。数据来源于开放同行评审平台OpenReview,该平台托管了顶级机器学习会议的评审过程。数据集重点关注围绕学术论文的同行评审生态系统,每条记录包括完整的评审相关信息:相关笔记(包含评审讨论、元评审、作者回复和社区反馈)、完整论文内容(Markdown格式)以及评审元数据(如页面统计、目录和文档结构分析)。评审数据捕获了完整的同行评审工作流程:来自多位评审的初始提交评审、作者反驳和回复轮次、领域主席的元评审、最终决定通知(接受/拒绝)以及发布后的讨论和社区评论。这使得数据集特别适用于:评审质量分析(研究同行评审质量和一致性的模式)、决策预测(基于论文内容和评审构建模型预测接受决定)、评审生成(训练模型生成建设性论文评审)、偏见检测(分析同行评审过程中的潜在偏见)以及科学话语分析(理解科学共识如何通过讨论形成)。数据集结构以JSON格式表示每篇论文及其相关评审数据,字段包括ID、标题、作者、摘要、年份、会议、相关笔记、PDF链接、来源链接、内容和内容元数据。
This dataset contains paper submissions and review data from the International Conference on Learning Representations (ICLR) for the years 2023, 2024, and 2025. The data is sourced from OpenReview, an open peer review platform that hosts the review process for top ML conferences. This dataset emphasizes the peer review ecosystem surrounding academic papers. Each record includes comprehensive review-related information: Related Notes (contains review discussions, meta-reviews, author responses, and community feedback from the OpenReview platform), Full Paper Content (complete paper text in Markdown format, enabling analysis of the relationship between paper content and review outcomes), and Review Metadata (structured metadata including page statistics, table of contents, and document structure analysis). The review data captures the full peer review workflow: initial submission reviews from multiple reviewers, author rebuttal and response rounds, meta-reviews from area chairs, final decision notifications (Accept/Reject), and post-publication discussions and community comments. This makes the dataset particularly valuable for: Review Quality Analysis (studying patterns in peer review quality and consistency), Decision Prediction (building models to predict acceptance decisions based on paper content and reviews), Review Generation (training models to generate constructive paper reviews), Bias Detection (analyzing potential biases in the peer review process), and Scientific Discourse Analysis (understanding how scientific consensus forms through discussion). The dataset structure represents each paper with its associated review data in JSON format, with fields including ID, title, authors, abstract, year, conference, related_notes, pdf_url, source_url, content, and content_meta.




