遇见数据集

A Multilingual Wikipedia Academic References Records

收藏
Zenodo2026-06-18 更新2026-06-21 收录
官方服务:

资源简介:

The dataset comprises two Apache Parquet files derived from 11.4 million citation templates extracted across six Wikipedia language editions and linked to the OpenAlex scholarly database. The first file, the Matched References file (multilang_wiki_ref_combined_openalex.parquet), contains all successfully resolved records, integrating original Wikimedia template metadata (such as title, authors, and academic identifier) with enriched OpenAlex fields including persistent identifiers (DOI, PMID), publication details, author and institution links, citation metrics, and related works. The second file, the Non-matched References file (multilang_wiki_ref_non_openalex.parquet) , preserves the remaining templates that could not be resolved to OpenAlex records; while it contains the same 7 core Wikimedia fields as the matched file, it specifically retains references that lack persistent identifiers, offering valuable data for analyzing citation practices, incomplete references, and non-scholarly sources that fall outside the scope of OpenAlex.

提供机构:
Zenodo
创建时间:
2026-06-18
二维码
社区交流群
二维码
科研交流群
商业服务