gsd-teacher-Indonesian
收藏资源简介:
该数据集包含4449个训练样本,采用结构化文本格式,总大小约为11.9MB。每个样本包含唯一标识符(id)、种子提示(seed_prompt)、语言(language)、模型(model)、消息列表(messages)和来源标识(source_id)字段。消息列表由一系列对话回合组成,每个回合包括角色(role)和内容(content)两个文本部分。数据集未提供背景、目的或内容描述,因此具体应用场景和任务类型无法确定。
This dataset contains 4,449 training samples, is stored in structured text format, and has a total size of approximately 11.9 MB. Each sample contains the following fields: unique identifier (id), seed prompt (seed_prompt), language (language), model (model), message list (messages), and source identifier (source_id). The message list is composed of multiple dialogue turns, where each turn includes two text components: role and content. No background, purpose, or content description is provided for this dataset, so its specific application scenarios and task types cannot be determined.
- 数据集名称:gsd-teacher-Indonesian
- 来源:Hugging Face Datasets
- 语言:印尼语
- 任务类型:对话数据,包含多轮消息(messages)
- 数据集规模:训练集共 5417 条样本
- 数据大小:下载大小约 12.28 MB,数据集总大小约 14.64 MB
- 数据字段:
id:字符串,样本唯一标识seed_prompt:字符串,种子提示词language:字符串,语言(印尼语)model:字符串,使用的模型messages:列表,包含对话消息,每条消息包含role(字符串,角色)和content(字符串,内容)source_id:字符串,来源标识
- 数据划分:仅包含
train划分(单训练集) - 配置文件:默认配置
default,数据文件路径为data/train-*





