S2ORC
收藏资源简介:
S2ORC数据集是从Semantic Scholar文献语料库中筛选出来的综合数据集,涵盖了医学、物理、生物、计算机科学等多个领域的论文。该数据集提供了论文的全文、注释、作者信息、引用作品和注释内容元素。数据集的创建过程包括从S2ORC中提取数据,并根据不同的特征(如引用位置、上下文类型等)生成诊断数据集。该数据集主要用于评估和分析引用推荐模型的性能,旨在解决引用推荐系统中的多样性和标准化问题。
The S2ORC dataset is a comprehensive curated dataset derived from the Semantic Scholar literature corpus, covering scholarly papers across multiple disciplines including medicine, physics, biology, computer science, and more. This dataset provides full texts of papers, annotations, author information, cited works, and annotated content elements. The construction of this dataset involves extracting data from S2ORC and generating diagnostic datasets based on diverse features such as citation positions and context types. This dataset is primarily utilized for evaluating and analyzing the performance of citation recommendation models, aiming to address the diversity and standardization issues present in citation recommendation systems.




