GuideMe
收藏资源简介:
GuideMe是首个支持多领域流视频任务指导的基准数据集,用于训练和评估多模态大语言模型的闭环交互任务指导能力,涵盖指令、反馈、错误检测和纠正等维度。数据集包含2,458个视频实例,总时长223.7小时,产生47,775个流式交互样本,覆盖烹饪、物体操作、日常生活指导和健身四个任务领域。视频长度从0.5到41.2分钟不等(平均5.5分钟,中位数3.6分钟)。数据集分为GuideMe-Train(1,985个视频,177.0小时)和GuideMe-Test(473个视频,46.7小时)两部分。每个交互样本涵盖四种指导类型之一:下一步指令、完成反馈、错误检测和纠正指导。
GuideMe is the first benchmark dataset supporting multi-domain streaming video task guidance, designed to train and evaluate the closed-loop interactive task guidance capabilities of multimodal large language models, covering dimensions such as instructions, feedback, error detection and correction. The dataset contains 2,458 video instances with a total duration of 223.7 hours, generating 47,775 streaming interactive samples, covering four task domains: cooking, object manipulation, daily life guidance, and fitness. The video lengths range from 0.5 to 41.2 minutes, with an average of 5.5 minutes and a median of 3.6 minutes. The dataset is split into two subsets: GuideMe-Train (1,985 videos, 177.0 hours) and GuideMe-Test (473 videos, 46.7 hours). Each interactive sample falls into one of four guidance types: next-step instructions, completion feedback, error detection, and correction guidance.
数据集概述:GuideMe
GuideMe 是首个面向流式视频的多领域任务指导与干预基准数据集,旨在评估多模态大语言模型(MLLMs)在闭环交互式任务指导场景中的表现,涵盖指令提供、反馈、错误检测与纠正等完整流程。
数据集规模与组成
- 总视频实例:2,458个
- 总时长:223.7小时
- 流式交互样本数:47,775个
- 视频时长范围:0.5分钟至41.2分钟(平均5.5分钟,中位数3.6分钟)
数据划分
| 子集 | 视频数量 | 时长(小时) | 说明 |
|---|---|---|---|
| GuideMe-Train | 1,985 | 177.0 | 用于任务特定适配或微调 |
| GuideMe-Test | 473 | 46.7 | 留出测试集,用于所有评估 |
任务领域
涵盖四个任务领域:
- 烹饪
- 物体操作
- 日常生活指导
- 健身
交互样本类型
每个交互样本涵盖以下四种指导类型之一:
- 下一步指令(Next-step instructions)
- 完成反馈(Completion feedback)
- 错误检测(Error detection)
- 纠正指导(Corrective guidance)
评估框架
包含三方面评估体系,共同衡量:
- 序列级对齐(Sequence-level alignment)
- 干预时机(Intervention timing)
- 内容质量(Content quality)
发布信息
- 相关代码与数据正在整理中,即将发布。
- 论文发表于 ECCV 2026,预印本可在 arXiv 获取:
- 论文地址:https://arxiv.org/abs/2607.02991
- 项目主页:https://fawnliu.github.io/project/guideme/




