Red Pajama V2是一个用于训练大型语言模型的开放数据集,包含超过1000亿个文本文档,其中30亿个文档带有质量信号,20亿个文档是去重后的唯一文档。数据集支持多种语言(如英语、德语、法语、西班牙语和意大利语),并提供了详细的下载和使用示例。此外,README还介绍了如何根据质量信号过滤数据集,并提供了质量注释的详细说明。
--- license: apache-2.0 --- # Dataset Summary RedPajama-Instruct-Data is curated from a diverse collection of NLP tasks from both [P3 (BigScience)](https://huggingface.co/datasets/bigscience/P3) and
--- license: llama2 language: - en --- # llama-instruct This dataset was used to finetune [Llama-2-7B-32K-Instruct](https://huggingface.co/togethercomputer/Llama-2-7B-32K-Instruct). We follow the di