amkyawdev/kyaw-mm-v1-dataset
收藏资源简介:
这是一个用于大型语言模型微调的缅甸语聊天机器人数据集,包含总计3,030,000个样本,分为训练集(1,024,000行)、验证集(1,003,000行)和测试集(1,003,000行),数据格式为Parquet,总大小约1.64 GB。每个样本包含一个消息数组,其中有系统提示(固定为英文You are a helpful assistant that speaks Myanmar fluently.)、用户输入(缅甸语或英语)和助理回复(缅甸语)。数据集涵盖五个类别:问候(နှုတ်ဆက်ခြင်း)、道歉(တောင်းပန်ခြင်း)、幽默(ဟာသပြောခြင်း)、解释(ရှင်းပြခြင်း)和一般(ယေဘုယျ),每个类别有6,000个样本。数据集由缅甸开发者Aung Myo Kyaw创建,旨在支持缅甸AI社区,许可证为Apache-2.0。
This is a Burmese (Myanmar) chatbot dataset for fine-tuning large language models, containing a total of 3,030,000 samples divided into training set (1,024,000 rows), validation set (1,003,000 rows), and test set (1,003,000 rows). The data is in Parquet format with a total size of approximately 1.64 GB. Each sample includes a messages array with a system prompt (fixed as You are a helpful assistant that speaks Myanmar fluently. in English), user input (in Burmese or English), and assistant response (in Burmese). The dataset covers five categories: Greeting (နှုတ်ဆက်ခြင်း), Apology (တောင်းပန်ခြင်း), Humor (ဟာသပြောခြင်း), Explanation (ရှင်းပြခြင်း), and General (ယေဘုယျ), with 6,000 samples per category. It was created by Myanmar developer Aung Myo Kyaw to support the Myanmar AI community, and is licensed under Apache-2.0.



