openbmb/factnet_factsynset
收藏资源简介:
FactSynset是FactNet的语义等价层,它将相似的FactStatements聚合为具有规范化值的统一语义类。它提供了语义等价事实的跨语言视图,支持跨语言障碍的推理。数据集包含parquet文件,其中包含多个关键字段,如synset_id(语义等价类的唯一标识符)、aggregation_key(聚合键)、member_statement_ids(该synset中的FactStatement ID列表)等。该数据集可用于跨语言事实检查、多语言知识图谱补全和语义推理等高级应用。数据集基于Wikidata和Wikipedia,采用CC BY-SA许可证。
FactSynset is the semantic equivalence layer of FactNet that aggregates similar FactStatements into unified semantic classes with normalized values. It provides a cross-lingual view of semantically equivalent facts, enabling reasoning across language barriers. The dataset contains parquet files with key fields such as synset_id (unique identifier for the semantic equivalence class), aggregation_key (aggregation key), member_statement_ids (list of FactStatement IDs in this synset), etc. The dataset enables advanced applications like cross-lingual fact checking, multilingual knowledge graph completion, and semantic reasoning. It is derived from Wikidata and Wikipedia and is available under the CC BY-SA license.




