S2AND
收藏资源简介:
S2AND是一个针对学术论文的作者姓名消歧统一基准数据集,由艾伦人工智能研究所创建。该数据集整合了八个先前独立的数据集,形成一个统一格式,并采用Semantic Scholar数据库中的丰富特征集。S2AND旨在通过提供一个全面的数据资源,帮助研究人员评估和比较不同的作者姓名消歧算法。数据集涵盖了广泛的学术领域和作者特征,如出版年份、论文数量等,以支持更公平和全面的算法评估。此外,S2AND还提供了一个开源的参考模型实现,以及详细的评估套件,使研究人员能够跟踪算法的全球性能和跨不同特征值的公平性。
S2AND is a unified benchmark dataset for academic paper author name disambiguation, developed by the Allen Institute for AI. This dataset integrates eight previously independent datasets into a unified format, and leverages a rich set of features sourced from the Semantic Scholar database. S2AND is designed to help researchers evaluate and compare different author name disambiguation algorithms by providing a comprehensive data resource. The dataset covers a broad spectrum of academic disciplines and author-related characteristics, such as publication year, number of published papers, and more, to support more fair and comprehensive algorithmic evaluations. In addition, S2AND also offers an open-source reference model implementation and a detailed evaluation suite, enabling researchers to track the global performance of algorithms and their fairness across different feature values.




