PAIR-IQ
收藏资源简介:
PAIR-IQ数据集由上海交通大学与上海创新研究院创建,旨在为大型语言模型生成的研究想法提供客观、人类对齐的评估基准。该数据集包含来自ICLR 2024、ICLR 2025和NeurIPS 2024的11,164篇论文,涵盖不同接收类别(如oral、spotlight、poster、reject)及12个研究主题,每篇论文包含标准化评分与结构化的想法表示。创建过程包括从OpenReview收集评分、执行均值偏移去偏以消除会议间偏差,并利用大语言模型将论文想法形式化为主要目标、核心突破、创新方法和实验设计四个组件。该数据集旨在解决当前研究想法评估手段碎片化、缺乏统一标准的问题,为大规模、可复现的对比评估提供黄金参考,推动科学想法的自动化评估。
The PAIR-IQ Dataset was created by Shanghai Jiao Tong University and Shanghai Institute of Innovation to provide an objective, human-aligned evaluation benchmark for research ideas generated by large language models (LLMs). This dataset includes 11,164 papers from ICLR 2024, ICLR 2025 and NeurIPS 2024, covering diverse acceptance categories such as oral, spotlight, poster and reject, as well as 12 research topics, with each paper containing standardized ratings and structured idea representations. The dataset construction process involves collecting review scores from OpenReview, applying mean shift debiasing to eliminate cross-conference biases, and leveraging large language models to formalize each paper’s research ideas into four components: main objectives, core breakthroughs, innovative methods and experimental designs. This dataset aims to address the current issues of fragmented research idea evaluation methods and lack of unified standards, providing a gold-standard reference for large-scale, reproducible comparative evaluations and promoting automated evaluation of scientific research ideas.





