factnet_factsense
收藏资源简介:
FactSense数据集是FactNet的语言层,提供从维基百科页面提取的多语言自然语言事实表达。每个FactSense实例代表一个以自然文本实现的事实陈述,并包含来源信息。数据集包含parquet文件,关键字段包括唯一标识符、关联事实陈述ID、语言代码、维基百科页面信息、匹配类型、包含事实提及的文本片段、匹配置信度、实体和属性ID、标签信息等。该数据集支持多语言事实核查、基于知识的生成和跨语言信息检索等应用。数据集基于维基百科文本,采用CC BY-SA许可。
The FactSense dataset is the language layer of FactNet, providing multilingual natural language factual expressions extracted from Wikipedia pages. Each FactSense instance represents a factual statement expressed in natural text and includes source information. The dataset contains Parquet files, with key fields including unique identifiers, associated factual statement IDs, language codes, Wikipedia page information, matching types, text segments containing factual mentions, matching confidence scores, entity and attribute IDs, label information, and more. This dataset supports applications such as multilingual fact-checking, knowledge-based generation, and cross-lingual information retrieval. The dataset is based on Wikipedia text and is licensed under CC BY-SA.




