Section-level annotations dataset for scientific papers (based on S2ORC)
收藏资源简介:
该数据集是基于Semantic Scholar开放研究语料库(S2ORC)构建的大规模科学论文修辞结构标注资源,由斯坦福大学和哥本哈根大学等机构的研究团队开发。数据集包含约1350万篇论文的章节级标注,覆盖数亿个文本片段,主要来源于STEM领域(特别是医学和生物学)的开放获取学术文献。创建过程采用高效的两阶段基于规则分类算法,结合精确匹配和正则表达式,对GROBID解析后的章节标题进行识别与标注,并通过标签传播和质量过滤确保数据可靠性。该资源旨在支持科学计量学和大规模计算研究,用于分析科学话语模式、学科特定写作规范以及科学交流的历时演变。
This dataset is a large-scale rhetorical structure annotation resource for scientific papers, built upon the Semantic Scholar Open Research Corpus (S2ORC) and developed by research teams from institutions such as Stanford University and the University of Copenhagen. It contains chapter-level annotations for approximately 13.5 million papers, covering hundreds of millions of text segments, and is primarily derived from open-access academic literature across STEM fields, with a particular focus on medicine and biology. The dataset's construction employs an efficient two-stage rule-based classification algorithm, which combines exact matching and regular expressions to identify and annotate section titles parsed by GROBID, and ensures data reliability through label propagation and quality filtering. This resource aims to support scientometrics and large-scale computational research for analyzing scientific discourse patterns, discipline-specific writing conventions, and the diachronic evolution of scientific communication.

- 1Large-scale dataset of automatically classified rhetorical sections in scientific papers斯坦福大学·教育学院; 哥本哈根大学·社会数据科学中心; 哥本哈根信息技术大学·网络、数据与社会研究中心; 人工智能先锋中心; 复杂性科学中心 · 2026年




