caine-dataset
收藏资源简介:
该数据集旨在用于训练Caine模型,以支持通用的创意写作任务。其内容主要来源于网络上的各类叙事性和百科类文本,具体包括同人小说(Fanfiction)、情色文学(Erotica)、维基百科页面(Wikipedia)以及类似性质的资料。数据来源广泛,涵盖了多个知名平台和社区,例如Wattpad、Literotica、Rtenzo、Creepypasta、Nightscribe和Fandom等。数据集规模在10万到100万样本之间,语言为英语,主要适用于文本生成任务,尤其侧重于故事创作、同人小说和特定类型叙事内容的生成。
This dataset is designed for training the Caine model to support general creative writing tasks. Its content primarily consists of narrative and encyclopedic texts sourced from the internet, specifically including fanfiction, erotica, Wikipedia pages, and similar materials. It draws from a wide range of well-known platforms and communities, such as Wattpad, Literotica, Rtenzo, Creepypasta, Nightscribe, and Fandom, among others. The dataset contains between 100,000 and 1,000,000 English samples, and is primarily intended for text generation tasks, with a particular focus on story creation, fanfiction, and the generation of narrative content across specific genres.
数据集概述
- 数据集名称:Caine Dataset
- 许可证:MIT
- 任务类别:文本生成
- 语言:英语
- 数据规模:100,000 至 1,000,000 条样本
- 标签:故事、rtenzo、同人小说、情色文学
数据内容
该数据集用于训练面向一般创意写作的 Caine 模型,包含以下类型的内容:
- 同人小说
- Wattpad 平台内容
- Literotica 平台内容
- Rtenzo 平台内容
- Creepypasta 恐怖故事
- Nightscribe 平台内容
- Wikipedia 页面
- Fandom 平台内容
- 其他类似内容




