synkrisnew2
收藏资源简介:
BelleGroup/train-1M-zh是一个用于大语言模型(LLM)监督微调(SFT)的中文指令数据集。该数据集基于BelleGroup公开的0.5M数据,通过self-instruct方法扩展生成了1,000,000条数据,旨在提供丰富的指令-输出对以提升模型遵循指令的能力。数据格式为JSONL,每条样本包含两个字段:instruction(用户指令)和output(预期输出)。数据集被划分为训练集(900,000条)和验证集(100,000条)。数据可能存在一定噪音,建议用户在使用前进行清洗。该数据集适用于中文大语言模型的指令微调、对话生成等任务。
BelleGroup/train-1M-zh is a Chinese instruction dataset for supervised fine-tuning (SFT) of large language models (LLMs). Based on BelleGroups publicly available 0.5M data, this dataset was expanded using the self-instruct method to generate 1,000,000 entries, aiming to provide rich instruction-output pairs to enhance the models ability to follow instructions. The data format is JSONL, with each sample containing two fields: instruction (user instruction) and output (expected output). The dataset is divided into a training set (900,000 entries) and a validation set (100,000 entries). The data may contain some noise, and it is recommended that users clean it before use. This dataset is suitable for tasks such as instruction fine-tuning and dialogue generation for Chinese large language models.
数据集概述
基本信息
- 数据集名称: synkrisnew2
- 许可证: Apache-2.0
说明
该数据集名为 synkrisnew2,采用 Apache-2.0 开源许可证发布。当前提供的详细信息有限,仅包含上述基本元数据,未提供关于数据规模、内容类型、语言、任务领域或具体用途的描述。




