shenlanmn/pile-uncopyrighted
收藏资源简介:
Pile Uncopyrighted是一个基于The Pile数据集的无版权版本,移除了所有受版权保护的内容,以响应作者要求大型语言模型停止使用其作品的呼吁。该数据集旨在支持用户训练未来大型语言模型时尊重作者权利并遵守版权法。清理方法包括删除Books3、BookCorpus2、OpenSubtitles、YTSubtitles和OWT2子集中的所有数据,这些子集在原始论文中未明确允许用于AI训练。计划未来创建更大数据集(如RedPajama)的无版权版本,但暂无具体时间表。
Pile Uncopyrighted is a copyright-free version of The Pile dataset, with all copyrighted content removed, created in response to authors demanding that LLMs stop using their works. This dataset encourages users to train future LLMs while respecting authors and abiding by copyright law. The cleaning methodology involves removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2 subsets, as these are the only datasets not explicitly allowed for AI training per the original paper. There are plans to create an uncopyrighted version of a larger dataset (e.g., RedPajama) with no estimated timeline.



