TVQA+
收藏资源简介:
我们提出了时空视频问答的任务,这需要智能系统同时检索相关时刻并检测引用的视觉概念(人和物体)来回答有关视频的自然语言问题。我们首先使用 310.8k 边界框来扩充 TVQA 数据集,将描绘的对象与问题和答案中的视觉概念联系起来。我们将此增强版本命名为 TVQA+。然后,我们提出了基于证据的时空回答器(STAGE),这是一个统一的框架,可以在空间和时间域中建立证据来回答有关视频的问题。综合实验和分析证明了我们框架的有效性以及我们 TVQA+ 数据集中的丰富注释如何有助于问答任务。作为一个副产品,通过执行这个联合任务,我们的模型能够产生更有洞察力的中间结果。
We introduce the task of Spatio-Temporal Video Question Answering (ST-VQA), which requires intelligent systems to retrieve relevant moments and detect referred visual concepts (people and objects) simultaneously to answer natural language questions about videos. We first augment the TVQA dataset with 310.8k bounding boxes, linking depicted objects to the visual concepts in questions and answers, and name this enhanced version TVQA+. We then propose the Evidence-based Spatio-Temporal Answerer (STAGE), a unified framework that grounds evidence in both spatial and temporal domains to answer questions about videos. Comprehensive experiments and analyses demonstrate the effectiveness of our framework and how the rich annotations in our TVQA+ dataset contribute to the question answering task. As a byproduct, by executing this joint task, our model is capable of generating more insightful intermediate results.




