WD4P: A link prediction benchmark for knowledge graph with multiple annotation patterns.
收藏资源简介:
WD4P (Wikidata with 4 Patterns) is a benchmark for link prediction on knowledge graphs containing base triple and three annotation patterns. In addition to base triples (s-p-o), the annotation patterns are: t-p-o pattern: describe annotations where a triple is in subject position, s-p-t pattern: describe annotations where a triple is in object position, and t-p-t pattern: describe annotations where triples are in both positions. To build WD4P, we use WD50K [1] since it is a curated hyper-relational KG, and FBHE [2], since it is a bi-level KG, and because both KGs are based on FB15K-237. We join WD50K and FBHE by taking advantage of that last fact. Similar to WD50K, WD4P uses entities and relations from Wikidata (rather than Freebase), as Wikidata is still maintained. The different patterns were obtained like follows. We obtained the base triples (and s-p-o pattern) from the 233,711 base triples of WD50K and the annotated triples of FBHE. We obtained the t-p-o pattern by reusing half of the 46,645 annotations of WD50K. To obtain the s-p-t, we reversed the other half of the annotations of WD50K by incorporating the appropriate symmetric relations. We obtain the t-p-t pattern by reusing as much of the 34,941 annotations of FBHE as possible. This method has for consequence that WD4P contains at most one annotation pattern per fact. That is because facts from WD50K contain only the t-p-o, or the s-p-t pattern once reversed. While facts from FBHE are only a singular t-p-t pattern. For evaluation purposes, we built two subsets of WD4P. WD4P-HR, a hyper-relational KG containing only s-p-o and t-p-o patterns of WD4P. WD4P-BL, a bi-level KG containing only s-p-o and t-p-t patterns. WD4P and its variants are randomly split into the same training set, validation set, and test set, respectively containing 80%, 10%, and 10% of the original information. As was done for WD50K, we eliminate test set leakages by removing facts from the train and validation sets that share their annotated triple with a fact of the test set. Then, to ensure that every term in the test set has been seen during training, we remove facts from the test set that contain entities and relations not present in the train or validation sets. [1] Mikhail Galkin, Priyansh Trivedi, Gaurav Maheshwari, Ricardo Usbeck, and Jens Lehmann. 2020. Message Passing for Hyper-Relational Knowledge Graphs. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Bonnie Webber, Trevor Cohn, Yulan He, and Yang Liu (Eds.). Association for Computational Linguistics, 7346–7359. doi:10.18653/v1/ 2020.emnlp- main.596 [2] Chanyoung Chung and Joyce Jiyoung Whang. 2023. Learning Representations of Bi-level Knowledge Graphs for Reasoning beyond Link Prediction. Proceedings of the AAAI Conference on Artificial Intelligence 37, 4 (June 2023), 4208–4216. doi:10.1609/aaai.v37i4.25538



