遇见数据集

A Comprehensive Dataset of Citations with Identifiers from English Wikipedia

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The dataset is composed of <strong>3 parts</strong>: 1. The dataset of 23.8 million citations from 35 different citation templates, out of which 3.14 million citations already contained identifiers, and approximately 2.15 million citations were equipped with identifiers from Crossref. This is under the filename: <strong>citations_from_wikipedia.zip</strong> 2. An example subset with the features for the classifier. This is under the filename: <strong>subset_of_citations_features.zip</strong> 3. Citations classified as a journal and their corresponding metadata/identifier extracted from Crossref to make the dataset more complete. This is under the filename: <strong>lookup_data.zip</strong>. This zip file contains a CSV file: <strong>lookup_table.gzip</strong> (a parquet file containing all citations classified as a journal) and a folder:<strong> metadata_extracted </strong>(a folder containing the metadata from CrossRef for all the citations mentioned in the table) <br> The data was parsed from the Wikipedia XML content dumps published in October 2018. The source code to extract and getting used to the pipeline can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki</strong> The taxonomy of the dataset in (1) can be found here: <strong>https://github.com/Harshdeep1996/cite-classifications-wiki/wiki/Taxonomy-of-the-parent-dataset</strong>

提供机构:
Zenodo
创建时间:
2020-01-14
二维码
社区交流群
二维码
科研交流群
商业服务