factnet_factstatements
收藏资源简介:
FactStatement数据集是FactNet知识图谱的基础层,这是一个跨语言、多层次的真实知识图谱。FactStatements是语言中立的原子事实单元,直接从Wikidata声明映射而来,构成了知识图谱的核心构建块。数据集包含parquet文件,关键字段包括:核心ID(core_id)、主体实体Wikidata QID(subject_qid)、属性Wikidata PID(property_pid)、原始值(value)、实体QID(value_qid,若非实体则为null)、标准化值(normalized_value)、标准化哈希值(claim_hash)、限定信息(qualifiers)、来源信息(references)、Wikidata等级(rank)和计算置信度(confidence)。该数据集设计为语言中立的事实表示,可通过FactSense层进行语言实现,并通过FactSynset层进行语义分组。数据集基于Wikidata构建,采用CC0许可协议。
The FactStatement dataset serves as the foundational layer of FactNet, a multilingual, multi-level factual knowledge graph. FactStatements are language-neutral atomic factual units directly mapped from Wikidata statements, forming the core building blocks of the knowledge graph. The dataset includes Parquet files, with key fields including: Core ID (core_id), Wikidata QID of the subject entity (subject_qid), Wikidata PID of the property (property_pid), raw value (value), entity QID (value_qid, null if not an entity), normalized value (normalized_value), normalized hash (claim_hash), qualifiers, references, Wikidata rank, and calculated confidence (confidence). This dataset is designed as a language-neutral factual representation, which can be linguistically realized via the FactSense layer and semantically grouped through the FactSynset layer. This dataset is built upon Wikidata and released under the CC0 license.




