遇见数据集

TVQA+

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

我们提出了时空视频问答的任务,这需要智能系统同时检索相关时刻并检测引用的视觉概念(人和物体)来回答有关视频的自然语言问题。我们首先使用 310.8k 边界框来扩充 TVQA 数据集,将描绘的对象与问题和答案中的视觉概念联系起来。我们将此增强版本命名为 TVQA+。然后,我们提出了基于证据的时空回答器(STAGE),这是一个统一的框架,可以在空间和时间域中建立证据来回答有关视频的问题。综合实验和分析证明了我们框架的有效性以及我们 TVQA+ 数据集中的丰富注释如何有助于问答任务。作为一个副产品,通过执行这个联合任务,我们的模型能够产生更有洞察力的中间结果。

We introduce the task of Spatio-Temporal Video Question Answering (ST-VQA), which requires intelligent systems to retrieve relevant moments and detect referred visual concepts (people and objects) simultaneously to answer natural language questions about videos. We first augment the TVQA dataset with 310.8k bounding boxes, linking depicted objects to the visual concepts in questions and answers, and name this enhanced version TVQA+. We then propose the Evidence-based Spatio-Temporal Answerer (STAGE), a unified framework that grounds evidence in both spatial and temporal domains to answer questions about videos. Comprehensive experiments and analyses demonstrate the effectiveness of our framework and how the rich annotations in our TVQA+ dataset contribute to the question answering task. As a byproduct, by executing this joint task, our model is capable of generating more insightful intermediate results.

提供机构:
OpenDataLab
创建时间:
2022-08-10
搜集汇总
数据集介绍
TVQA+ 数据集图片
背景与挑战
背景概述
TVQA+是一个增强的时空视频问答数据集,在TVQA基础上扩充了31.08万个边界框注释,以关联视觉概念与问题答案。该数据集由北卡罗来纳大学教堂山分校于2019年发布,旨在支持智能系统通过检索时刻和检测视觉对象来回答视频相关问题。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务