StoryVideoQA
收藏资源简介:
StoryVideoQA是由武汉大学研究团队构建的当前规模最大的深度视频理解数据集,旨在解决复杂故事情节的长视频理解难题。该数据集涵盖393.2小时的长篇故事视频,包含3部电视剧集和78部高评分电影,共生成超过36.3万个高质量问答对,平均视频时长分别为1635秒和7878秒。数据集通过升级的StoryMindv2多智能体协作框架自动生成,采用监督引导生成机制和多评审投票策略确保数据质量与主题平衡性。该数据集主要应用于深度视频理解领域,支持对长程角色关联、多层次故事元素和复杂推理能力的系统性评估,为视频问答模型的发展提供重要基准。
StoryVideoQA, constructed by the research team at Wuhan University, is the largest-scale deep video understanding dataset to date, designed to address the challenges of long-form video understanding with complex narrative plots. It encompasses 393.2 hours of long-form story videos, including 3 TV drama series and 78 highly-rated films, with a total of over 363,000 high-quality question-answer pairs. The average durations of the video samples for the TV series and film subsets are 1635 seconds and 7878 seconds, respectively. The dataset is automatically generated via the upgraded StoryMindv2 multi-agent collaboration framework, which adopts supervised-guided generation mechanisms and a multi-reviewer voting strategy to ensure data quality and topic balance. Primarily applied in the field of deep video understanding, this dataset supports systematic evaluations of long-range character relationships, multi-level story elements, and complex reasoning capabilities, serving as a critical benchmark for the development of video question answering models.

- 1StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset武汉大学·计算机学院; 国家多媒体软件工程技术研究中心; 湖北省多媒体与网络通信工程重点实验室 · 2026年



