遇见数据集

OpenAlex Author Name Disambiguation V3 Data - Disambiguation Model

收藏
Zenodo2023-07-31 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

5 Separate files used in the OpenAlex (https://openalex.org) V3 Author Name Disambiguation Model Creation: ORCID_hard_negative_pairs: Pairs of ORCIDs where either the full name, family name, or given name are a match and would therefore be more difficult to disambiguate. Disambiguator_all_possible_training_data: Dataset created which contains all possible features for modeling and all possible samples of data. Eventually, this was split into train/val/test and also processed more to create a better balance of positive to negative samples for our purposes. Disambiguator_final_train_data: Final data which the disambiguator was trained on. Disambiguator_final_val_data: Data which was used to test the model during training to optimize the features/hyperparameters chosen. Disambiguator_final_test_data: Final dataset which gave model performance indication after all hyperparameters were tuned and features were chosen. More details can be found at https://github.com/ourresearch/openalex-name-disambiguation

提供机构:
Zenodo
创建时间:
2023-07-31
二维码
社区交流群
二维码
科研交流群
商业服务