遇见数据集

GenWiki: A Dataset of 1.3 Million Content-Sharing Text and Graphs for Unsupervised Graph-to-Text Generation

收藏
DataCite Commons2025-05-08 更新2024-07-13 收录
官方服务:

资源简介:

Paper: "GenWiki: A Dataset of 1.3 Million Content-Sharing Text and Graphs for Unsupervised Graph-to-Text Generation" (COLING 2020) by Zhijing Jin, Qipeng Guo, Xipeng Qiu, and Zheng Zhang. (https://aclanthology.org/2020.coling-main.217/) <br> <br> Abstract: Data collection for the knowledge graph-to-text generation is expensive. As a result, research on unsupervised models has emerged as an active field recently. However, most unsupervised models have to use non-parallel versions of existing small supervised datasets, which largely constrain their potential. In this paper, we propose a large-scale, general-domain dataset, GenWiki. Our unsupervised dataset has 1.3M text and graph examples, respectively. With a human-annotated test set, we provide this new benchmark dataset for future research on unsupervised text generation from knowledge graphs.

论文:《GenWiki:面向无监督图到文本生成的130万内容共享文本与图谱数据集》(COLING 2020),作者为金志静、郭启鹏、邱锡鹏、张正。链接:https://aclanthology.org/2020.coling-main.217/ 摘要:知识图谱到文本生成(Knowledge Graph-to-Text Generation)任务的数据采集成本高昂,因此无监督模型相关研究近年来成为活跃研究领域。然而当前多数无监督模型仅能依托现有小规模监督数据集的非并行版本开展实验,这在很大程度上限制了其性能潜力。本文提出大规模通用领域数据集GenWiki,该无监督数据集分别包含130万条文本样本与图谱样本。我们配套构建了人工标注测试集,为后续基于知识图谱的无监督文本生成研究提供全新的基准数据集。

提供机构:
Edmond
创建时间:
2024-01-02
二维码
社区交流群
二维码
科研交流群
商业服务