遇见数据集

EmmaLeonhart/loka

收藏
Hugging Face2026-05-16 更新2026-05-31 收录
官方服务:

资源简介:

Loka是一个神经符号世界模型数据集,包含RDF-star三元存储语料库和训练好的Transformer检查点。语料库基于Wikidata数据,通过预处理将实体和属性的标识符(QIDs/PIDs)替换为英文标签,并过滤了非语义数据类型(如外部ID、URL等),以保留实体链接、字符串、数值和时间等语义内容。数据集包含训练输入的三元组文件、BPE词汇表、分词器配置以及模型生成的RDF-star推理三元组(附带溯源信息)。检查点采用44.5M参数的角色感知掩码Transformer架构,支持通过SPARQL查询生成三元组的溯源上下文。数据集旨在支持可审计、可查询的神经符号知识图谱研究,适用于特征提取和文本生成任务。

Loka is a neuro-symbolic world model dataset that contains an RDF-star triple storage corpus and trained Transformer checkpoints. The corpus is based on Wikidata data, where entity and property identifiers (QIDs/PIDs) are replaced with English labels through preprocessing, and non-semantic data types such as external IDs and URLs are filtered out to retain semantic content including entity links, strings, numerical values and timestamps. The dataset includes training input triple files, the BPE vocabulary, tokenizer configurations, and RDF-star inference triples generated by the model with attached provenance information. The checkpoints adopt a 44.5M-parameter role-aware masked Transformer architecture, which supports generating provenance contexts of triples via SPARQL queries. This dataset is designed to support auditable and queryable neuro-symbolic knowledge graph research, and is suitable for feature extraction and text generation tasks.

提供机构:
EmmaLeonhart
二维码
社区交流群
二维码
科研交流群
商业服务