遇见数据集

A Comprehensive Dataset of Classified Citations with Identifiers from English Wikipedia (2023)

收藏
Zenodo2024-04-28 更新2026-05-29 收录
数据链接:
官方服务:

资源简介:

This is a dataset of 40.664.485 citations extracted from English Wikipedia February 2023 dump (https://dumps.wikimedia.org/enwiki/20230220/). Version 1: en_citations.zip is a dataset of extracted citations Version 2: en_final.zip is the same dataset with classified citations augmented with identifiers The fields are as follows: type_of_citation - Wikipedia template type used to define the citation, e.g., 'cite journal', 'cite news', etc. page_title - title of the Wikipedia article from which the citation was extracted. Title - source title, e.g., title of the book, newspaper article, etc. URL - link to the source, e.g., webpage where news article was published, description of the book at the publisher's website, online library webpage, etc. tld - top link domain extracted from the URL, e.g., 'bbc' for https://www.bbc.co.uk/... Authors - list of article or book authors, if available. ID_list - list of publication identifiers mentioned in the citation, e.g., DOI, ISBN, etc. citations - citation text as used in Wikipedia code actual_label - 'book', 'journal', 'news', or 'other' label assigned based on the analysis of citation identifiers or top link domain. acquired_ID_list - identifiers located via Google Books and Crossref APIs for citations which are likely to refer to books or journals, i.e., defined using 'cite book', 'cite journal', 'cite encyclopedia', and 'cite proceedings' templates. The total number of news: 9.926.598 The total number of books: 2.994.601 The total number of journals: 2.052.172 Augmented with IDs via lookup 929.601 (out of 2.445.913 book, journal, encyclopedia, and proceedings template citations not classified as books or journals via given identifiers). The source code to extract citations can be found here: <strong>https://github.com/albatros13/wikicite. </strong> The code is a fork of the earlier project on Wikipedia citation extraction: https://github.com/Harshdeep1996/cite-classifications-wiki.

本数据集包含从2023年2月的英文维基百科快照(https://dumps.wikimedia.org/enwiki/20230220/)中提取的40,664,485条引用记录。版本1:en_citations.zip为提取得到的原始引用数据集;版本2:en_final.zip为在版本1基础上,为分类后的引用补充了标识符的增强版数据集。 字段说明如下: 1. type_of_citation:用于定义该引用的维基百科模板类型,例如`cite journal`、`cite news`等。 2. page_title:提取该引用的维基百科条目标题。 3. Title:来源文献标题,例如书籍、新闻报道的标题等。 4. URL:指向来源资源的链接,例如新闻报道发布网页、出版商官网的书籍介绍页面、在线图书馆网页等。 5. tld:从URL中提取的顶级链接域名,例如对于`https://www.bbc.co.uk/...`,其tld为`bbc`。 6. Authors:若有可用信息,则为文章或书籍的作者列表。 7. ID_list:引用中提及的出版物标识符列表,例如数字对象唯一标识符(DOI, Digital Object Identifier)、国际标准书号(ISBN, International Standard Book Number)等。 8. citations:维基百科代码中使用的原始引用文本。 9. actual_label:基于对引用标识符或顶级链接域名的分析所分配的标签,可选值为`book`(书籍)、`journal`(期刊)、`news`(新闻)或`other`(其他)。 10. acquired_ID_list:通过Google Books和Crossref接口为符合条件的引用补充获取的标识符,这类引用指使用`cite book`、`cite journal`、`cite encyclopedia`及`cite proceedings`模板,且未通过给定标识符归类为书籍或期刊的引用。 统计数据如下: - 新闻类引用总数:9,926,598条 - 书籍类引用总数:2,994,601条 - 期刊类引用总数:2,052,172条 通过标识符查找共为929,601条符合条件的引用补充了标识符,该类引用为未通过给定标识符归类为书籍或期刊的书籍、期刊、百科全书及会议论文集模板引用,总计2,445,913条。 用于提取引用的源代码可在此处获取:<strong>https://github.com/albatros13/wikicite</strong>。该代码为此前维基百科引用提取项目的分支项目,原项目地址为:https://github.com/Harshdeep1996/cite-classifications-wiki。

提供机构:
Zenodo
创建时间:
2023-07-03
二维码
社区交流群
二维码
科研交流群
商业服务