Problematic ROR-Affilliation Names in Crossref 2024 Dump
收藏资源简介:
Overview This dataset contains 326 entries corresponding to DOIs and related Affiliation Name and ROR IDs for which a mismatch between the name and ID was detected in Crossref data, April 2024 dump. It is a manually checked dataset of wrong affiliation names from a list of automatically pre-selected candidates. It may be used as a benchmark for matching algorithms working with affiliation data in Crossref. Entries stemming from some particular issues (3 ROR IDs with multiple issues) were not included, as they were considered less useful for the dataset as a benchmark of what wrong matches may look like (see the "Scripts and Analytics" contents for details.Note: the entries in the dataset represent entries and ROR-Affiliation Name pairs with issues (sometimes referred to as "false matches"). The pipeline focused on precision over recall, so it is not comprehensive and it is likely that there are other problematic entries in the 2024 dump not listed here. Source datasets The following CC0 datasets were used as source for this dataset: April 2024 Public Data File from CrossRef (http://doi.org/10.13003/849J5WP), downloaded via torrent ROR Release v1.59 (https://doi.org/10.5281/zenodo.14728473), downloaded manually via web browser Wikidata, queried via QLever (https://qlever.cs.uni-freiburg.de/wikidata), full Wikidata dump from https://dumps.wikimedia.org/wikidatawiki/entities (latest-all.ttl.bz2 and latest-lexemes.ttl.bz2, version 29.01.2025) Column meanings On the .tsv dataset (main), the column names are: DOI - The Crossref DOI for the work Affiliation_Name - An affiliation name string listed for some author of the work (DOI) ROR_ID - The ROR ID provided by the publisher corresponding to this Affiliation Name for this DOI ROR_Display - The display name for this ROR ID via the ROR Release v1.59 Status - "manually curated false match" for all; this is just a sanity check for data reusers, reinforcing these entries are manually curated to be wrongThe .xlsx file contains extra information and some notes done during the curation process. Scripts and analytics Scripts and analytics for the baseline matching pipeline are available (as of March 2025) at https://github.com/lubianat/crossref_interview.Manual curation was done in Google Sheets, available (as of March 2025) at https://docs.google.com/spreadsheets/d/1XX_v5sI_EYHtRUp69s5LjITJD7v2dp4JqdFvLolG23U/edit?gid=1978804245#gid=1978804245 with parts of the process live streamed at https://www.youtube.com/watch?v=-Jum8E3_cQs .
数据集概览 本数据集包含326条条目,对应2024年4月Crossref数据转储中检测到机构名称与ROR ID(Research Organization Registry ID)不匹配的数字对象标识符(DOI,Digital Object Identifier)、相关机构名称及ROR ID。该数据集是从自动预筛选的候选集中手动核验得到的错误机构名称数据集,可作为Crossref机构数据匹配算法的基准测试集。 部分因特定问题(3个存在多重问题的ROR ID)的条目未被纳入,原因是此类条目作为错误匹配样例的基准参考价值较低(详细信息参见“脚本与分析”章节)。注意:本数据集的条目代表存在问题的条目与ROR-机构名称对(有时也称为“错误匹配”)。该流程优先考量精确率而非召回率,因此数据集并不全面,2024年转储数据中可能还存在其他未被收录的问题条目。 源数据集 本数据集的来源数据集均采用CC0协议: 1. Crossref 2024年4月公开数据文件(http://doi.org/10.13003/849J5WP),通过种子文件下载 2. ROR 版本v1.59(https://doi.org/10.5281/zenodo.14728473),通过浏览器手动下载 3. Wikidata,通过QLever(https://qlever.cs.uni-freiburg.de/wikidata)查询,数据源为https://dumps.wikimedia.org/wikidatawiki/entities 上的完整Wikidata转储文件(latest-all.ttl.bz2与latest-lexemes.ttl.bz2,2025年1月29日版本) 列含义说明 本数据集的主文件为制表符分隔值(TSV,Tab-Separated Values)格式,各列含义如下: - DOI:对应作品的Crossref数字对象标识符 - Affiliation_Name:该作品(DOI)的某位作者所标注的机构名称字符串 - ROR_ID:出版方为该DOI对应的此机构名称提供的ROR ID - ROR_Display:通过ROR v1.59版本获取的该ROR ID对应的展示名称 - Status:全部条目均标注为“人工核验错误匹配”;该字段用于提醒数据复用者,确认本数据集所有条目均经人工核验确认为错误匹配。 附带的XLSX文件包含额外信息及人工核验过程中记录的部分注释。 脚本与分析 截至2025年3月,基线匹配流程的脚本与分析文件可在https://github.com/lubianat/crossref_interview 获取。人工核验工作在Google Sheets中完成,相关表格(截至2025年3月)可通过https://docs.google.com/spreadsheets/d/1XX_v5sI_EYHtRUp69s5LjITJD7v2dp4JqdFvLolG23U/edit?gid=1978804245#gid=1978804245 访问,部分核验过程的直播录像可在https://www.youtube.com/watch?v=-Jum8E3_cQs 观看。



