遇见数据集

Large-scale dataset of automatically classified rhetorical sections in scientific papers

收藏
Zenodo2026-07-03 更新2026-08-02 收录
官方服务:

资源简介:

We present a dataset of section-level annotations for millions of scientific papers from the Semantic Scholar Open Research Corpus (S2ORC). Using a rule-based classification algorithm, we identified and labeled major sections across 15.6 million papers. The dataset covers primarily STEM disciplines, with strong representation in medicine and biology.This dataset enables large-scale computational studies of scientific discourse, writing patterns, and the adoption of AI tools in different parts of the research communication process.The dataset is organized into a single file "classified.parquet" containing the corresponding section labels for every header in each paper. Each record includes the paper identifier from S2ORC, the header-section label, and metadata indicating the header's position within the document. We do not include the full text content of each header in our dataset, as this would duplicate content already available in S2ORC. Users can retrieve the corresponding text by matching our paper identifiers and header position metadata to the original S2ORC corpus, ensuring compatibility while minimizing redundancy. The file "aggregated.parquet" contains a reduced summary of "classified.parquet". Each record includes the paper identifier from S2ORC, the section label and metadata containing field classification of the paper, the number of characters in the section, the number of characters in the paper, the number of characters in the paper when removing trailing sections like references.Additional details on the data construction process, validation, usage notes and recommended filtering are described in the accompanying manuscript. You can find the preprint at

提供机构:
Zenodo
创建时间:
2026-06-29
二维码
社区交流群
二维码
科研交流群
商业服务