eval-fsr-a1-stack-jest-swe-r525-traces
收藏资源简介:
该数据集是一个结构化的对话实验数据集,专门用于记录和分析AI模型在特定任务上的交互表现。核心包含多轮对话记录,每轮对话均标注了发言内容和发言角色。此外,数据集提供了丰富的实验元数据,包括对话所使用的智能体、模型名称、模型提供商、实验日期、任务类型、实验轮次、运行标识符、试验名称、任务执行结果、验证器的输出评估以及数据追踪来源。数据集包含3,667个训练样本,总数据量约为529MB,适用于AI对话模型评估、多轮对话分析、智能体行为研究和实验过程追踪等任务。
This dataset is a structured dialogue experiment dataset specifically designed for recording and analyzing the interaction performance of AI models on specific tasks. The core consists of multi-turn dialogue records, with each turn annotated for content and role. Additionally, the dataset provides rich experimental metadata, including the agent used in the dialogue, model name, model provider, experiment date, task type, episode number, run identifier, trial name, task execution result, verifier output evaluation, and trace source. It contains 3,667 training samples, with a total data volume of approximately 529MB, and is suitable for tasks such as AI dialogue model evaluation, multi-turn dialogue analysis, agent behavior research, and experiment process tracking.
- 数据集名称:eval-fsr-a1-stack-jest-swe-r525-traces
- 提供机构:LAION
- 数据集描述:该数据集包含对话记录、代理信息、模型信息、任务执行细节及验证结果,适用于评估和追踪AI系统在多轮交互中的表现。
- 数据字段:
conversations:对话列表,每条对话包含content(内容,字符串)和role(角色,字符串)。agent:代理标识(字符串)。model:模型标识(字符串)。model_provider:模型提供商(字符串)。date:日期(字符串)。task:任务名称(字符串)。episode:回合编号(字符串)。run_id:运行ID(字符串)。trial_name:试验名称(字符串)。result:结果(字符串)。verifier_output:验证器输出(字符串)。trace_source:追踪来源(字符串)。
- 数据规模:
- 总下载大小:439.8 MB
- 总数据集大小:529.1 MB
- 训练集样本数:3,667 条
- 数据划分:
- 训练集(train):包含全部数据
- 数据文件格式:数据文件位于
data/train-*,默认配置为default。




