A Multilingual Wikipedia Academic References Records
收藏资源简介:
The dataset comprises two Apache Parquet files derived from 11.4 million citation templates extracted across six Wikipedia language editions and linked to the OpenAlex scholarly database. The first file, the Matched References file (multilang_wiki_ref_combined_openalex.parquet), contains all successfully resolved records, integrating original Wikimedia template metadata (such as title, authors, and academic identifier) with enriched OpenAlex fields including persistent identifiers (DOI, PMID), publication details, author and institution links, citation metrics, and related works. The second file, the Non-matched References file (multilang_wiki_ref_non_openalex.parquet) , preserves the remaining templates that could not be resolved to OpenAlex records; while it contains the same 7 core Wikimedia fields as the matched file, it specifically retains references that lack persistent identifiers, offering valuable data for analyzing citation practices, incomplete references, and non-scholarly sources that fall outside the scope of OpenAlex.



