We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpu
Zero LLM数据集提供了从多个公共数据集中提取的经过清理和标准化的文本语料库,专门为书籍《从零开始深度学习❻——LLM篇》准备。每个语料库都组织在自己的目录中(如codebot/、storybot/、webbot/),以便文本数据、BPE编码数据和合并规则分组在一起。这使得数据集易于理解、导航和复现。数据集包括Python代码、TinyStories V2和OpenWebText样本,提供原始
--- task_categories: - text-generation language: - en pretty_name: Red Pajama 1T --- ### Getting Started The dataset consists of 2084 jsonl files. You can download the dataset using HuggingFace: ```p