DCAgent2/swebench_verified_random_100_folders_g1_top8_31600_32b_20260430_163909
收藏资源简介:
该数据集是一个包含多轮对话记录的结构化数据集,用于分析和评估对话代理或模型在特定任务中的表现。数据集特征包括对话内容(conversations,包含角色和内容)、代理类型(agent)、模型名称(model)、模型提供商(model_provider)、日期(date)、任务类型(task)、剧集(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)和验证器输出(verifier_output)。数据仅包含训练分割(train),有300个示例,总大小约44.4 MB。数据集可能适用于自然语言处理、对话系统评估或任务完成度分析等应用场景。
This is a structured dataset containing multi-turn conversation records, designed for analyzing and evaluating the performance of dialogue agents or models on specific tasks. The dataset's features include conversation content (conversations, including roles and their respective utterances), agent type (agent), model name (model), model provider (model_provider), date (date), task type (task), episode, run ID (run_id), trial name (trial_name), result (result), and verifier output (verifier_output). The dataset only includes the training split (train), with 300 examples and a total size of approximately 44.4 MB. This dataset may be applicable to scenarios such as natural language processing, dialogue system evaluation, and task completion analysis.




