tonebridge-metrics
收藏资源简介:
Chinese-Llama-3-8B-Instruct数据集是一个专门用于训练中文指令跟随模型的高质量指令数据集,包含约200万条指令-输出对。数据来源多样,整合了来自Firefly、Belle、MOSS、ShareGPT等多个公开数据集的指令数据,以及通过自研方法构建的高质量中文指令数据。数据集经过严格的清洗和去重处理,确保数据质量,并覆盖广泛的主题领域。每条数据包含instruction(指令)和output(输出)两个字段,形成标准的指令-响应格式,旨在支持中文指令理解与生成任务的模型训练,适用于构建能够理解和执行自然语言指令的人工智能助手。数据集采用Apache-2.0许可证发布,可通过HuggingFace datasets库直接加载使用。
The Chinese-Llama-3-8B-Instruct dataset is a high-quality instruction dataset specifically designed for training Chinese instruction-following models, containing approximately 2 million instruction-output pairs. It features diverse data sources, integrating instruction data from multiple public datasets such as Firefly, Belle, MOSS, and ShareGPT, as well as high-quality Chinese instruction data constructed through proprietary methods. The dataset undergoes rigorous cleaning and deduplication to ensure data quality and covers a wide range of topics. Each data entry includes two fields: instruction and output, forming a standard instruction-response format. It aims to support model training for Chinese instruction understanding and generation tasks, suitable for building AI assistants capable of understanding and executing natural language instructions. The dataset is released under the Apache-2.0 license and can be directly loaded using the HuggingFace datasets library.
基于提供的页面地址和README文件内容,该数据集只有以下信息:
- 许可证:Apache-2.0
除此之外,README文件中没有提供任何关于数据集的具体描述、内容、用途、规模、结构或其它元数据信息。该数据集详情页面的内容极为有限。




