gsd-teacher-Swahili
收藏资源简介:
该数据集是一个结构化的对话数据集,包含4510个训练样本。每个样本代表一次模型生成的对话交互,核心数据是一个对话消息列表(messages),其中每条消息包含发言者角色(role)和内容(content)。对话的发起基于一个文本提示(seed_prompt)。数据集还记录了生成该对话所使用的模型(model)、对话的语言(language)、样本的唯一标识符(id)以及来源标识(source_id)。该数据集适用于对话系统构建、大语言模型指令遵循能力评估、多语言对话分析等自然语言处理任务。
This dataset is a structured conversational dataset containing 4510 training samples. Each sample represents a model-generated dialogue interaction, with core data being a list of dialogue messages (messages), where each message includes the speaker role (role) and content (content). The dialogue is initiated based on a text prompt (seed_prompt). The dataset also records the model used to generate the dialogue (model), the language of the dialogue (language), a unique identifier for the sample (id), and a source identifier (source_id). It is suitable for natural language processing tasks such as dialogue system construction, evaluation of large language model instruction-following capabilities, and multilingual dialogue analysis.
- 数据集名称:gsd-teacher-Swahili
- 数据集页面:https://huggingface.co/datasets/ljvmiranda921/gsd-teacher-Swahili
- 语言:斯瓦希里语(Swahili)
- 数据集规模:
- 下载大小:约 11 MB
- 数据集大小:约 13.38 MB
- 训练集样本数量:5446 条
- 数据特征:
id:标识符(字符串类型)seed_prompt:种子提示(字符串类型)language:语言(字符串类型)model:模型(字符串类型)messages:消息列表(包含角色role和内容content,均为字符串类型)source_id:来源ID(字符串类型)
- 数据划分:仅包含
train划分 - 配置:默认配置
default,数据文件路径为data/train-*




