遇见数据集

Proposed threshold-based and rule-based approaches to detecting duplicates in bibliographic database

收藏
Zenodo2024-06-24 更新2024-06-25 收录
数据链接:
官方服务:

资源简介:

Bibliographic databases are used to measure the performance of researchers, universities and research institutions. Thus, high data quality is required and data duplication is avoided. One of the weaknesses of the threshold-based approach in duplication detection is the low accuracy level. Therefore, another approach is required to improve duplication detection. This study proposes a method that combines threshold-based and rule-based approaches to perform duplication detection. These two approaches are implemented in the comparison stage. The cosine similarity function is used to create weight vectors from the features. Then, the comparison operator is used to determine whether the pair of records are grouped as duplication or not. Three research databases: Web of Science (WoS), Scopus, and Google Scholar (GS) on the Science and Technology Index (SINTA) database are investigated. Rule 4 and Rule 5 provide the best performance. For WoS dataset, the accuracy, precision, recall, and F1-measure values were 100.00%. For Scopus dataset, the accuracy and precision values were 100.00%, recall: 98.00%, and the F1-measure value is 98.00%. For GS dataset, the accuracy value was 100.00%, precision: 99.00%, recall: 97.00%, and the F1-measure value is 98.00%. The proposed method is potential tool for accurate detection on duplication records in publication databases.

文献书目数据库(Bibliographic databases)常用于评估研究人员、高校及科研机构的学术表现,因此此类数据库对数据质量要求极高,且需避免数据重复。在重复数据检测任务中,基于阈值(threshold-based)的方法存在准确率较低的短板,因此亟需开发更优的重复数据检测方案。本研究提出一种融合基于阈值与基于规则(rule-based)两种方法的重复数据检测方案,该方案在比对阶段集成了这两类方法:研究采用余弦相似度函数(cosine similarity function)从特征维度生成权重向量,随后通过比对算子判断记录对是否应被归类为重复数据。本研究针对收录于科技索引(Science and Technology Index, SINTA)数据库中的Web of Science(WoS)、Scopus及Google Scholar(GS)三大学术数据集展开实验。实验结果显示,规则4与规则5展现出最优的实验性能;在WoS数据集上,准确率、精确率、召回率及F1值均达到100.00%;在Scopus数据集上,准确率与精确率均为100.00%,召回率为98.00%,F1值为98.00%;在GS数据集上,准确率为100.00%,精确率为99.00%,召回率为97.00%,F1值为98.00%。所提方案可作为学术出版数据库中精准检测重复记录的有效工具。

提供机构:
M. Miftakul Amin
创建时间:
2024-06-24
二维码
社区交流群
二维码
科研交流群
商业服务