Wutaghost/LLMscore-ICLR-OpenReview
收藏资源简介:
LLMscore-ICLR-OpenReview数据集是论文《Position: Peer Review Should Be Calibrated via LLM Scoring》的原始发布数据集,专门用于同行评审分析。该数据集旨在研究论文评审理由、数值评分、LLM衍生的锚定分数和评审分数残差在科学同行评审中的相互作用。它不是一个通用论文语料库,使用时需考虑其特定的方法论背景。数据集包含ICLR 2023、2024和2025年的OpenReview链接数据,组织为论文元数据、标准化评审记录、匿名优缺点理由项、锚定分数、偏差值和提取的论文文本。数据集不包含论文PDF,但提供了openreview_url和pdf_url以访问原始内容。提取的论文文本以UTF-8 .txt文件形式存储在每年的tar.gz压缩包中。数据集还包含模式信息、数据集级摘要统计和安全说明,支持同行评审校准分析、评审理由和评分一致性分析、论文级和评审级残差/偏差分析,以及相关论文实验的复现或扩展。
The LLMscore-ICLR-OpenReview dataset is the released original dataset for the paper Position: Peer Review Should Be Calibrated via LLM Scoring. Its concrete purpose is peer review analysis: the dataset is meant for studying how paper-review rationales, numeric ratings, LLM-derived anchor scores, and review-score residuals interact in scientific peer review. It is not a generic paper corpus, and it should be used with that specific methodological context in mind. The package organizes OpenReview-linked ICLR 2023, 2024, and 2025 data, containing paper metadata, normalized review records, anonymized pro/con reason items, anchor scores, bias values, and extracted paper text. The package intentionally does not include paper PDFs, but provides openreview_url and pdf_url to access the original content. Extracted paper text is stored in yearly tar.gz bundles as UTF-8 .txt files. The dataset also includes schema information, dataset-level summary statistics, and safety notes, supporting peer review calibration analysis, reviewer rationale and score-consistency analysis, paper-level and review-level residual/bias analysis, and reproducing or extending experiments from the associated paper.





