遇见数据集

CHEN-AND: A Labeled Dataset for Chinese and English Joint Author Name Disambiguation

收藏
Zenodo2024-10-05 更新2026-05-28 收录
官方服务:

资源简介:

Abstract: Author name disambiguation (AND) is an important problem in literature databases and is even more prominent in the cross-language (database) context. In order to eliminate such ambiguity, extensive research has been conducted in the academic community. However, the existing research mainly focuses on monolingual literature, with less attention paid to author disambiguation in cross-language environments. In this regard, this study focuses on a typical cross-language author disambiguation task - the disambiguation of joint Chinese and English authors in academic literature. We first propose an automated dataset construction method for Chinese-English literature joint AND using online open resources, with this method, we create a dataset named CHEN-AND for joint Chinese and English author disambiguation. Then we propose a merging-first-then-disambiguation (MFTD) based author disambiguation framework and evaluate several variants of this method on the test dataset. The experimental results show that the P-F1 and B3-F1 accuracies of the best-performing method among the MFTD variants are only 83.66% and 88.46%, which is lower than the accuracies of mainstream monolingual author disambiguation methods, indicating that there is still much room for performance improvement. For details on how to create this dataset, please refer to this repository on my GitHub page.

提供机构:
Zenodo
创建时间:
2023-12-02
二维码
社区交流群
二维码
科研交流群
商业服务