遇见数据集

Manually collected Citation Context Datasets from moderately cited Biomedical Publications

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

This dataset consists of citation contexts that were collected from one hundred (100) moderately cited biomedical articles that were sampled from the MEDLINE database. The 100 sampled articles received 7317 citations (maximum=179, minimum=31, average=73.17) from which 11,228 citation contexts were extracted. The citation contexts were manually identified as the span of texts around the citation marker that describes the contribution that was referenced from the cited articles. Unlike other studies that identified citation contexts using a predetermined window of texts before and after the citation marker, citation contexts in this dataset were identified by reading the texts around the citation marker to understand the context in which a cited publication was referenced. Citation contexts data that are collected using a predetermined window of words/paragraphs are prone to errors(overrepresentation or underrepresentation). The first sheet in the excel files contains bibliographic details of the cited article, while the second and third sheets contain bibliographic details of the cited articles and the citation contexts, respectively. Four information types were collected on the third sheet containing the citation contexts: No, title, number of mentions, and citation contexts. "No" refers to a citing article's unique serial number that also corresponds to the serial number of the citing article on the second sheet of the excel file. "Title" refers to the title of the citing article, "number of mentions" connotes the number of times the cited article is referenced in the citing article, and "citation context" is the text that represents the contribution of the cited article. This dataset is useful for citation identification, classification, and weighting studies. This dataset was collected as part of the first author's doctoral thesis.

本数据集包含从MEDLINE(MEDLINE)数据库中抽取的100篇被适度引用的生物医学文献的引用上下文(citation contexts)。这100篇抽样文献总计被引用7317次(最高179次,最低31次,均值73.17次),从中共提取出11228条引用上下文(citation contexts)。本数据集的引用上下文均通过人工识别,即定位引用标记(citation marker)周围用于描述被引文献所贡献内容的文本片段。与其他研究通过预先设定引用标记前后的文本窗口来识别引用上下文的方法不同,本数据集的引用上下文是通过阅读引用标记周围的文本,理解被引文献被引用的具体语境后完成识别的。通过预先设定词/段落窗口的方式采集的引用上下文数据,易出现过度表征或表征不足等误差。Excel文件的首个工作表包含被引文献的书目详细信息,第二和第三工作表则分别对应被引文献的书目详情与引用上下文数据。在包含引用上下文的第三工作表中,共记录了四类信息:序号(No)、文献标题、提及次数以及引用上下文。“序号(No)”指代引用文献的唯一序列号,该序列号与Excel文件第二工作表中对应引用文献的序列号一致。“文献标题”指代引用文献的标题;“提及次数”表示被引文献在该引用文献中被引用的总次数;“引用上下文”则为描述被引文献所贡献内容的文本片段。本数据集可应用于引用识别、分类以及权重计算等相关研究。本数据集的采集工作作为第一作者博士学位论文的一部分完成。

创建时间:
2021-05-17
二维码
社区交流群
二维码
科研交流群
商业服务