gsd-smith-Swahili
收藏资源简介:
该数据集包含用于对话生成或代理行为分析任务的结构化样本。每个样本由以下核心字段构成:唯一标识符(id)、初始对话种子提示(seed_prompt)、语言标识(language)、生成回复所使用的模型名称(model)、一个多轮对话消息列表(messages,其中每条消息包含角色和内容),以及一个记录代理决策或行动轨迹的JSON结构(agent_trace)。此外,样本还包含一个来源标识符(source_id)。数据集以训练集(train)形式提供,共包含717个样本,数据总量约为15.6MB。该数据集适用于研究对话系统、多轮交互、代理行为建模或作为相关机器学习任务的训练与评估数据。
This dataset contains structured samples intended for dialogue generation or agent behavior analysis tasks. Each sample comprises the following core fields: a unique identifier (id), an initial dialogue seed prompt (seed_prompt), a language identifier (language), the name of the model used to generate responses (model), a multi-turn dialogue message list (messages, where each entry contains a role and corresponding content), and a JSON structure recording the agent's decision-making or action trajectory (agent_trace). Additionally, each sample includes a source identifier (source_id). The dataset is provided as a training split (train), containing a total of 717 samples with an approximate total data size of 15.6 MB. This dataset is suitable for research on dialogue systems, multi-turn interactions, agent behavior modeling, or serving as training and evaluation data for related machine learning tasks.




