Gold Standard Benchmark Dataset for Bibliometric Reference Matching Evaluation
收藏资源简介:
This dataset provides the gold standard benchmark used in the paper: Aria, M., D'Aniello, L., Spano, M. (2026). "A Multi-Phase Reference Matching Algorithm for Bibliometric Analysis: Design, Implementation, and Evaluation." Scientometrics (submitted). The dataset consists of 1,064 journal articles retrieved from Web of Science using the query TI=("bibliometrics" OR "science mapping"), restricted to English-language journal articles with complete metadata. For each article, the following fields are included: author list (full names), article title, journal name (full title and ISO 4 abbreviation), publication year, volume, issue, start page, end page, and DOI. Each article was converted into a Scopus-format reference string, with journal names randomly assigned as either full title or ISO 4 abbreviation (approximately 50% each), introducing realistic heterogeneity in journal name representation. The resulting reference strings, each associated with a unique article identifier, constitute the gold standard against which the 17 perturbation scenarios described in the paper are evaluated. The dataset supports the reproducibility of the synthetic benchmark experiments presented in the paper and can be reused for evaluating other reference matching or record linkage algorithms in bibliometric contexts.



