ANetQA
收藏资源简介:
ANetQA是一个大规模的视频问答基准数据集,旨在支持对未剪辑视频的细粒度组合推理。该数据集由杭州电子科技大学计算机学院创建,包含1340万个平衡的问答对。数据集中的问答对自动从标注的视频场景图中生成,反映了视频的细粒度语义、时空场景图和多样化的问答模板。ANetQA适用于评估视频问答模型的多种推理能力,如对象识别、关系理解和属性分析,旨在推动视频与语言学习领域的研究。
ANetQA is a large-scale video question answering (QA) benchmark dataset designed to support fine-grained compositional reasoning over untrimmed videos. Developed by the School of Computer Science, Hangzhou Dianzi University, this dataset includes 13.4 million balanced question-answer pairs. All QA pairs within the dataset are automatically generated from annotated video scene graphs, which capture the fine-grained semantic information, spatiotemporal scene graphs, and diverse QA templates of the corresponding videos. ANetQA can be utilized to evaluate diverse reasoning abilities of video QA models, including object recognition, relational understanding, and attribute analysis, with the objective of advancing research in the video-and-language learning domain.




