gsd-smith-Yoruba
收藏资源简介:
该数据集包含1672个训练样本,总大小约47.4MB,用于支持对话系统研究、多轮对话生成、智能体行为分析和跨语言对话任务。其结构包括核心字段:唯一标识符(id)、种子提示(seed_prompt)、语言类型(language)、模型信息(model)、多轮对话消息(messages)、智能体轨迹(agent_trace)、来源标识(source_id)和研究早期停止标记(research_early_stopping)。其中,messages字段是结构化列表,每条消息包含角色(role)和内容(content);agent_trace字段存储JSON格式的列表数据,特别适合需要追踪对话历史和智能体决策过程的研究场景。
This dataset contains 1672 training samples with a total size of approximately 47.4MB, designed to support research in dialogue systems, multi-turn conversation generation, agent behavior analysis, and cross-lingual dialogue tasks. Its structure includes core fields: unique identifier (id), seed prompt (seed_prompt), language type (language), model information (model), multi-turn conversation messages (messages), agent trace (agent_trace), source identifier (source_id), and research early stopping marker (research_early_stopping). The messages field is a structured list where each message contains a role and content, while the agent_trace field stores JSON-formatted list data, making it particularly suitable for research scenarios that require tracking conversation history and agent decision-making processes.




