ScienceMeter
收藏资源简介:
我们从10个科学领域中每个领域检索了1000篇期刊或会议论文,使用Semantic Scholar API。对于每篇论文,我们还收集了其引用论文,形成了我们的原始语料库。我们过滤掉缺乏引用信息或摘要的论文,然后根据给定模型的知识截止日期和论文的出版日期重新分组剩余的论文。这一过程产生了5,148个三元组(先前论文、新论文、未来论文)。对于每篇论文,我们合成了一个支持声明(一个独特的支持科学主张)和一个反驳声明(一个相关但不支持的科学主张)。生成的数据集可在filtered_with_claims文件夹中找到。
We retrieved 1000 journal or conference papers from each of 10 scientific fields using the Semantic Scholar API. For each paper, we also collected its cited papers, forming our original corpus. We filtered out papers lacking citation information or abstracts, and then regrouped the remaining papers based on the knowledge cutoff date of the given model and the publication date of the papers. This process produced 5,148 triplets (previous paper, new paper, future paper). For each paper, we synthesized a supporting statement (a unique supporting scientific argument) and a rebuttal statement (a related but non-supporting scientific argument). The generated dataset is located in the 'filtered_with_claims' folder.
ScienceMeter数据集概述
数据集简介
- 名称:ScienceMeter
- 用途:追踪语言模型中科学知识的更新情况
- 状态:开发中(代码和文档可能不完整或会变动)
数据来源
- 通过Semantic Scholar API获取10个科学领域的期刊或会议论文
- 每个领域收集1,000篇论文及其引用论文,形成原始语料库(raw corpus)
数据处理
- 过滤标准:
- 剔除缺少引用信息或摘要的论文
- 重组方法:
- 根据模型的知识截止日期和论文发表日期进行分组
- 最终数据:
- 5,148个三元组(先前论文,新论文,未来论文)
数据增强
- 为每篇论文生成:
- 1个SUPPORT声明(独特的支持性科学主张)
- 1个REFUTE声明(相关但不支持的科学主张)
数据可用性
- 最终数据集存放路径:
filtered_with_claims文件夹
引用信息
latex @article{wang2025sciencemeter, title={ScienceMeter: Tracking Scientific Knowledge Updates in Language Models}, author={Wang, Yike and Feng, Shangbin and Tsvetkov, Yulia and Hajishirzi, Hannaneh}, journal={arXiv preprint arXiv:2505.24302}, year={2025} }




