vn-1-dataset
收藏资源简介:
VN-1训练数据集是一个用于基于Llama 3.2-3B-Instruct模型微调VN-1模型的越南语指令微调数据集。该数据集包含73个高质量的越南语对话样本,数据格式为结构化的聊天消息,包含系统、用户和助手角色。对话主题广泛,涵盖文化、历史、知识、健康、技术和日常生活等多个领域。该数据集适用于文本生成任务,特别是越南语聊天模型的指令微调与对齐。
The VN-1 training dataset is a Vietnamese instruction fine-tuning dataset used for fine-tuning the VN-1 model based on the Llama 3.2-3B-Instruct model. It contains 73 high-quality Vietnamese dialogue samples in a structured chat message format, including system, user, and assistant roles. The dialogue topics are broad, covering various fields such as culture, history, knowledge, health, technology, and daily life. This dataset is suitable for text generation tasks, particularly for instruction fine-tuning and alignment of Vietnamese chat models.
VN-1 训练数据集
- 语言: 越南语 (vi)
- 许可证: Apache-2.0
- 任务类别: 文本生成
- 标签: 越南语、聊天、指令微调、VN-1
- 数据集规模: 少于 1000 个样本 (n<1K)
数据集详情
- 样本数量: 73 个高质量越南语对话样本
- 数据格式: 聊天消息(包含 system / user / assistant 角色)
- 主题范围: 文化、历史、知识、健康、科技、生活
用途说明
该数据集用于基于 Llama 3.2-3B-Instruct 模型对 VN-1 模型进行微调。




