遇见数据集

ray0rf1re/AO3-2020

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

AO3-2020是一个基于Archive of Our Own(ao3.org)平台的英文同人小说语料库,包含截至2020年的创意写作和长篇小说内容。数据集总文档数为15,993,173,总字符数约21,816,825,882,估计token数约5,454,206,470,原始源大小为502 GB(SQLite格式),压缩为Parquet格式(使用ZSTD level 3压缩)。语言为英语,许可证为CC BY-NC 4.0(仅限非商业研究使用)。数据集适用于文本生成、填充掩码、特征提取和语言建模等NLP任务,并包含多个子集配置,如完整语料库、令牌上限子集(用于受控规模实验)、分数子集(确定性均匀采样)以及增强子集(使用OpenHermes 2.5/ChatML指令模板格式化,适用于直接SFT训练)。数据模式包括text字段(UTF-8文本)和保留的原始SQLite元数据列。数据集还包含一个安全对齐子集(adult_18plus_100K),用于偏好对训练中的拒绝侧,以提高模型安全性。

AO3-2020 is an English fanfiction corpus based on the Archive of Our Own (ao3.org) platform, containing creative writing and long-form literature up to 2020. The dataset includes 15,993,173 total documents, approximately 21,816,825,882 total characters, and an estimated 5,454,206,470 tokens. The raw source size is 502 GB (in SQLite format), compressed into Parquet files with ZSTD level 3 compression. The language is English, and it is licensed under CC BY-NC 4.0 (for non-commercial research use only). It is suitable for NLP tasks such as text generation, fill-mask, feature extraction, and language modeling. The dataset offers multiple subset configurations, including the full corpus, token-capped subsets (for controlled-scale experiments), fraction subsets (deterministic uniform sampling), and augmented subsets (formatted with the OpenHermes 2.5/ChatML instruction template for direct SFT training). The data schema features a text field (UTF-8 text) and preserved original SQLite metadata columns. It also includes a safety/alignment subset (adult_18plus_100K) intended for use as the rejected side in preference pair training to enhance model safety.

提供机构:
ray0rf1re
二维码
社区交流群
二维码
科研交流群
商业服务