gsd-teacher-Japanese
收藏资源简介:
该数据集包含5406个训练样本,数据总量约16.2MB。每个样本由唯一标识符(id)、种子提示(seed_prompt)、语言(language)、生成模型(model)、消息序列(messages)以及源ID(source_id)构成。消息序列为列表结构,每条消息包含角色(role)和内容(content)两个子字段,表明数据以多轮对话或交互形式组织。基于字段特征,数据集可能适用于语言模型微调、对话系统开发、提示工程优化或跨模型响应分析等任务。
This dataset contains 5,406 training samples with a total data volume of approximately 16.2 MB. Each sample consists of a unique identifier (id), a seed prompt, a language field, a generation model (model), a message sequence (messages), and a source ID (source_id). The message sequence is structured as a list, where each message contains two sub-fields: role and content, indicating that the data is organized in the form of multi-turn dialogues or interactions. Based on the characteristics of its fields, this dataset may be applicable to tasks such as language model fine-tuning, dialogue system development, prompt engineering optimization, and cross-model response analysis.
数据集概述
- 数据集名称:gsd-teacher-Japanese
- 数据集地址:https://huggingface.co/datasets/ljvmiranda921/gsd-teacher-Japanese
数据集详情
- 任务类型:对话生成(基于教师模型的日语对话数据)
- 语言:日语
- 数据集大小:
- 下载大小:13,990,546 字节
- 数据集总大小:16,197,908 字节
数据划分
| 划分名称 | 样本数量 | 字节数 |
|---|---|---|
| train | 5,406 | 16,197,908 |
特征字段
| 字段名 | 类型 | 描述 |
|---|---|---|
| id | string | 样本唯一标识符 |
| seed_prompt | string | 初始提示文本 |
| language | string | 语言(日语) |
| model | string | 生成消息所使用的模型 |
| messages | list | 包含角色(role)和内容(content)的消息列表 |
| source_id | string | 来源标识符 |
数据文件
- 配置文件名称:default
- 数据文件路径:https://huggingface.co/datasets/ljvmiranda921/gsd-teacher-Japanese/tree/main/data/train-*(通配符表示多个文件)




