遇见数据集

pg19

收藏
OpenCSG2024-07-19 更新2026-01-19 收录
官方服务:

资源简介:

PG-19主要用于语言建模基准测试,它包含从Project Gutenberg图书库中提取的1919年之前出版的一系列书籍,以及书籍标题和出版日期等元数据。数据集规模超过Billion Word基准的两倍,文档平均长度是WikiText长程语言建模基准的20倍。书籍被划分为训练集、验证集和测试集,元数据存储在包含book_id、short_book_title和publication_date的metadata.csv文件中。PG-19不限制词汇量大小,而是以开放词汇基准的形式发布数据,仅对文本进行了少量处理,例如删除样板许可文本,并将Ofcom指定的冒犯性歧视词映射到占位符。该数据集支持长程语言模型的基准测试,也可用于预训练其他需要长程推理的自然语言处理任务。PG-19采用Apache 2.0许可。

PG-19 is primarily designed as a benchmark for language modeling. It consists of a corpus of books published before 1919, extracted from the Project Gutenberg library, alongside metadata such as book titles and publication dates. The total size of this dataset is more than twice that of the Billion Word benchmark, and the average length of its documents is 20 times that of the WikiText long-range language modeling benchmark. The books are partitioned into training, validation, and test sets. The metadata is stored in the metadata.csv file, which contains fields including book_id, short_book_title, and publication_date. PG-19 does not enforce a fixed vocabulary size; instead, it is released as an open-vocabulary benchmark, with only minimal text preprocessing applied—such as removing boilerplate license text and mapping offensive discriminatory terms specified by Ofcom to placeholder tokens. This dataset supports benchmarking long-range language models, and can also be utilized for pre-training other natural language processing tasks that require long-range reasoning capabilities. PG-19 is licensed under the Apache 2.0 License.

提供机构:
AIWizards
创建时间:
2024-07-19
搜集汇总
数据集介绍
pg19 数据集图片
背景与挑战
背景概述
PG-19是一个用于长程语言建模基准测试的数据集,包含从Project Gutenberg图书库提取的1919年之前出版的书籍及其元数据(如标题和出版日期)。该数据集规模庞大,是Billion Word基准的两倍以上,文档平均长度是WikiText基准的20倍,采用开放词汇设计,仅进行了最小化文本处理,适用于长程推理任务的预训练和评估。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务