CUSP
收藏资源简介:
CUSP(Cutoff-conditioned Unseen Scientific Progress)是一个用于评估人工智能系统在历史知识截止条件下预测科学进展能力的基准数据集。其核心目标是测试模型的前瞻性推理能力,即给定一个模型的训练知识截止日期,该模型能否预测在该日期之后实际发生的真实科学发现。数据集聚焦于评估科学预见性,而非对已知事实的推理。数据内容来源于六种权威科学出版物,涵盖八个核心研究领域,包括生物学、医学、化学、物理学、环境科学、材料科学、细胞生物学和人工智能。数据的时间覆盖范围从2023年3月到2026年2月。数据集规模在1,000到10,000个样本之间。每个数据样本对应一篇真实发表的科学论文,并以JSON Lines格式存储,包含丰富的元数据和多种任务格式,如样本唯一标识符、论文来源、高层研究领域、细粒度研究子领域、论文标题、摘要、数字对象标识符、真实发表日期以及模型知识截止日期。数据集设计了五种不同的任务变体来多角度评估模型:二元问题、扰动二元问题、多项选择题、自由回答题和日期预测。该数据集适用于问答、文本分类和文本生成等任务,特别用于科学预测、时间推理和知识截止感知的模型评估,并采用双轨框架进行严格评估,数据经过多阶段验证以确保质量。
CUSP (Cutoff-conditioned Unseen Scientific Progress) is a benchmark dataset for evaluating the ability of artificial intelligence (AI) systems to predict scientific progress under knowledge cutoff constraints. Its core objective is to test the model's prospective reasoning capability: given the knowledge cutoff date of the model's training data, can it predict the real scientific discoveries that actually occurred after that date? This dataset focuses on evaluating scientific foresight rather than reasoning about known facts. The dataset content is sourced from six authoritative scientific publications, covering eight core research disciplines including biology, medicine, chemistry, physics, environmental science, materials science, cell biology, and artificial intelligence. The temporal coverage of the dataset spans from March 2023 to February 2026. The dataset contains between 1,000 and 10,000 samples. Each data sample corresponds to a real published scientific paper, stored in JSON Lines format, and includes rich metadata and multiple task-related fields, such as unique sample identifier, paper source, high-level research domain, fine-grained research sub-domain, paper title, abstract, Digital Object Identifier (DOI), actual publication date, and model knowledge cutoff date. The dataset designs five distinct task variants to evaluate the model from multiple perspectives: binary classification tasks, perturbed binary classification tasks, multiple-choice questions, free-response questions, and date prediction tasks. This dataset is applicable to tasks such as question answering, text classification, and text generation, and is specifically used for model evaluation in scientific prediction, temporal reasoning, and knowledge cutoff awareness. It adopts a dual-track framework for rigorous evaluation, and the data has undergone multi-stage validation to ensure quality.




