遇见数据集

REASONS: REtrieval and Automated citationS Of scieNtific Sentences

收藏
Zenodo2026-01-16 更新2026-05-26 收录
官方服务:

资源简介:

The dataset spans 12 scientific domains, comprising 10 in computer science and 2 in biology. In total, REASONS contains 19,904 source papers, of which 3,944 satisfy IEEE formatting constraints, yielding 12,723 sentence-level citation instances. The dataset exhibits a pronounced long-tail structure. Computer Vision is the most heavily represented domain, contributing 5,488 papers and 3,437 cited sentences, while specialized domains such as Biomolecules and Quantum Computing contribute substantially fewer papers 119 and 421, respectively) and citation instances 27 and 456. Mid-sized domains such as Robotics, Graphics, Information Retrieval, Artificial Intelligence, and Natural Language Processing collectively form the head of the distribution. This structure is intentional rather than a limitation. It enables controlled comparison between well-represented domains with relatively standardized terminology and specialized domains that exhibit sparse coverage, domain-specific vocabulary, and irregular citation practices. These contrasts reveal when models abstain versus when they confidently misattribute.

本数据集涵盖12个科学领域,其中计算机科学领域10个,生物学领域2个。总体而言,REASONS数据集共包含19904篇源文献,其中3944篇符合IEEE格式规范,最终生成12723个句子级引用实例。该数据集呈现出显著的长尾分布结构。计算机视觉(Computer Vision)是占比最高的领域,涵盖5488篇源文献与3437个被引句子;而生物分子学(Biomolecules)与量子计算(Quantum Computing)等细分领域的源文献数量分别仅为119篇和421篇,对应引用实例分别为27个和456个,占比极低。机器人学、图形学、信息检索、人工智能(Artificial Intelligence)以及自然语言处理(Natural Language Processing)等中等规模领域则共同构成了该分布的头部区间。该长尾结构是刻意设计的结果,而非数据集本身的局限。这种结构支持对两类领域进行可控对比:一类是术语相对标准化、样本充足的主流领域,另一类则是覆盖稀疏、带有专属词汇体系且引用行为不规范的细分领域。通过这类对比,可以清晰观察到模型何时会选择弃权,何时又会自信地做出错误归因。

提供机构:
Zenodo
创建时间:
2026-01-16
二维码
社区交流群
二维码
科研交流群
商业服务