SciCiteVal
收藏资源简介:
SciCiteVal数据集专为引文验证任务设计,包含人工标注的引文标签,分为“正确”、“错误”和“无关”三类。对于标注为“错误”的引文,进一步定义了五个子类别以描述不准确的性质。每个数据样本由引用论文中的**引文上下文**和被引论文中支持标签的**证据段落**组成。数据集包含1,034条引文,分布在机器学习与生物学领域的科学论文中,其中302条正确引文、302条错误引文和430条无关引文。正确和错误引文改编自QASA数据集,无关引文则从实际论文中提取。数据集包含四列数据:“引文上下文”、“被引内容”、“标签”和“扭曲类别”,适用于文本分类任务。数据标注过程包括对QASA数据的验证转换(正确引文)、系统性扭曲(错误引文)以及跨领域手动收集(无关引文)。数据集以TSV格式提供,采用CC-BY-4.0许可协议。
The SciCiteVal dataset is specifically designed for the citation verification task, containing manually annotated citation labels divided into three categories: "Correct", "Incorrect", and "Irrelevant". For citations labeled as "Incorrect", five subcategories are further defined to describe the nature of the inaccuracy. Each data sample consists of the citation context from the citing paper and the evidence passage from the cited paper that supports the assigned label. The dataset contains 1,034 citations from scientific papers in the fields of machine learning and biology, including 302 correct citations, 302 incorrect citations, and 430 irrelevant citations. Correct and incorrect citations are adapted from the QASA dataset, while irrelevant citations are extracted from real-world scientific papers. The dataset includes four columns: "citation context", "cited content", "label", and "distortion category", and is applicable to text classification tasks. The data annotation process includes validation and conversion of QASA data (for correct citations), systematic distortion (for incorrect citations), and cross-domain manual collection (for irrelevant citations). The dataset is provided in TSV format and is licensed under CC-BY-4.0.



