遇见数据集

EmmaLeonhart/normalized-wikidata

收藏
Hugging Face2026-05-14 更新2026-05-31 收录
官方服务:

资源简介:

Normalized Wikidata是一个预处理后的文本形式Wikidata数据集,专门优化用于训练语言模型或知识图谱世界模型。其核心目标是提供一个干净的语料库,突出Wikidata三元组的语义内容,同时去除原始数据中占主导地位的目录和标识符杂乱信息。数据集以每行一个三元组的格式呈现,使用制表符分隔,所有位置(主语、谓语、宾语)均为英文标签,其中实体标签来自源转储,属性标签来自一个包含7,312个手动解析的Wikidata属性的精选缓存。在预处理过程中,移除了非语义属性(如外部ID、URL、数学公式等),并规范化了时间、数量值和多语言文本值,以解决原始Wikidata中的已知问题,如目录爆炸、属性标签损坏和数据类型后缀泄漏。该数据集旨在支持高效的世界模型训练,提供清晰的语义信号。

Normalized Wikidata is a preprocessed textual Wikidata dataset specifically optimized for training language models or knowledge graph world models. Its core objective is to provide a clean corpus that highlights the semantic content of Wikidata triples, while eliminating the pervasive clutter of catalog-style entries and identifiers present in the raw dataset. The dataset is formatted as one triple per line, with fields separated by tab characters; all three positions (subject, predicate, object) use English labels, where entity labels are sourced from the original Wikidata dump, and predicate labels are derived from a curated cache of 7,312 manually parsed Wikidata properties. During preprocessing, non-semantic properties (e.g., external IDs, URLs, mathematical formulas, etc.) are removed, and temporal, quantitative, and multilingual textual values are normalized to resolve known issues in the raw Wikidata dataset, including catalog explosion, corrupted property labels, and data type suffix leakage. This dataset is intended to support efficient world model training by delivering clear semantic signals.

提供机构:
EmmaLeonhart
二维码
社区交流群
二维码
科研交流群
商业服务