用于结构功能识别的大规模数据集
收藏资源简介:
该数据集由南京理工大学信息管理系的研究人员创建,旨在通过自动识别科学论文中的结构功能,以提升科学论文的摘要质量。数据集由从arXiv和PubMed收集的原始文章组成,这些文章的章节标题被标准化为IMRaD格式,以便进行结构功能识别。通过训练一个分类器来自动识别章节中的关键结构组件,如背景、方法、结果、讨论等。最后,使用Longformer模型来捕捉丰富的上下文信息,生成科学论文的摘要。
This dataset was developed by researchers from the Department of Information Management, Nanjing University of Science and Technology, with the goal of improving the quality of scientific paper abstracts by automatically recognizing structural functions within scientific papers. The dataset includes original articles collected from arXiv and PubMed, whose section titles have been standardized to the IMRaD format to enable structural function recognition. A classifier is trained to automatically identify key structural components in sections, such as background, methods, results, discussions and other categories. Finally, the Longformer model is employed to capture rich contextual information and generate abstracts for scientific papers.




