WhoIsWho
收藏资源简介:
WhoIsWho数据集是由清华大学开发的一个大规模学术名称消歧基准,包含超过100万个文档。该数据集通过交互式标注过程构建,涉及10多名专业标注者,历时约24个月。数据集内容丰富,包括文档的标题、作者名、组织、关键词、摘要、出版年份和会议/期刊信息。WhoIsWho数据集旨在解决在线学术系统中名称消歧这一基本问题,特别是在研究论文数量不断增长的背景下,提高算法在处理大规模和高质量数据集上的有效性。该数据集的应用领域广泛,包括但不限于学术搜索平台的优化、学术合作网络的构建以及学术评价系统的改进。
The WhoIsWho dataset is a large-scale academic name disambiguation benchmark developed by Tsinghua University, containing over one million documents. It was constructed through an interactive annotation process, involving more than 10 professional annotators and taking approximately 24 months to complete. The dataset is rich in content, including the title, author names, affiliations, keywords, abstracts, publication years, and conference/journal information of the documents. The WhoIsWho dataset aims to address the fundamental problem of name disambiguation in online academic systems, particularly to improve the effectiveness of algorithms when processing large-scale and high-quality datasets against the backdrop of the ever-growing number of research papers. The dataset has a wide range of application scenarios, including but not limited to the optimization of academic search platforms, the construction of academic collaboration networks, and the improvement of academic evaluation systems.

- 1A framework for constructing a huge name disambiguation dataset: algorithms, visualization and human collaboration清华大学 · 2020年



