encyclopedic-completion-corpus
收藏资源简介:
百科全书补全语料库是一个用于文本生成和问答任务的大规模英文数据集,专注于事实性知识的多词补全。它提供基于年份条件的长短语补全对,涵盖科学常数、人物(如创始人)、地理信息、事件日期和算术运算等领域。每个数据样本包括提示和对应的补全短语,短语长度通常为8-20个词,超越单一实体标记。数据集还标注了知识截止年份(表示事实在该年份前有效)、类别和来源。数据规模为每个文件约5500万行,按年份组织(如2013年)。类别分布近似为:科学常数占30%,事件日期占26%,人物创始人占22%,地理占13%,算术占9%。数据来源于公共百科全书式时间线、传记、地理资料、离线CC0风格静态表格/公式、合成算术,以及可选的Wikidata SPARQL查询(CC0许可)。该数据集适用于训练或评估模型在百科全书式事实补全、知识问答和文本生成任务中的性能。
The Encyclopedic Completion Corpus is a large-scale English dataset for text generation and question-answering tasks, focusing on multi-word completion of factual knowledge. It provides year-conditioned long-phrase completion pairs covering areas such as scientific constants, people (e.g., founders), geographic information, event dates, and arithmetic operations. Each data sample includes a prompt and a corresponding completion phrase, typically 8-20 words long, going beyond single-entity tokens. The dataset also annotates the knowledge cutoff year (indicating when the fact was valid), category, and source. The data scale is approximately 55 million lines per file, organized by year (e.g., 2013). The category distribution is roughly: scientific constants 30%, event dates 26%, people founders 22%, geography 13%, arithmetic 9%. Data sources include public encyclopedic timelines, biographies, geographic materials, offline CC0-style static tables/formulas, synthetic arithmetic, and optional Wikidata SPARQL queries (CC0 license). This dataset is suitable for training or evaluating models in encyclopedic fact completion, knowledge-based QA, and text generation tasks.





