REASONS: REtrieval and Automated citationS Of scieNtific Sentences
收藏资源简介:
The dataset spans 12 scientific domains, comprising 10 in computer science and 2 in biology. In total, REASONS contains 19,904 source papers, of which 3,944 satisfy IEEE formatting constraints, yielding 12,723 sentence-level citation instances. The dataset exhibits a pronounced long-tail structure. Computer Vision is the most heavily represented domain, contributing 5,488 papers and 3,437 cited sentences, while specialized domains such as Biomolecules and Quantum Computing contribute substantially fewer papers 119 and 421, respectively) and citation instances 27 and 456. Mid-sized domains such as Robotics, Graphics, Information Retrieval, Artificial Intelligence, and Natural Language Processing collectively form the head of the distribution. This structure is intentional rather than a limitation. It enables controlled comparison between well-represented domains with relatively standardized terminology and specialized domains that exhibit sparse coverage, domain-specific vocabulary, and irregular citation practices. These contrasts reveal when models abstain versus when they confidently misattribute.



