FortisAVQA
收藏资源简介:
FortisAVQA是一个专为评估音频视觉问答模型鲁棒性而设计的数据集。该数据集通过两个阶段构建:首先,对MUSIC-AVQA公开数据集的测试部分问题进行重写,以增强多样性;其次,引入基于答案分布的问题分布偏移,以实现精细的鲁棒性评估。该数据集保留了MUSIC-AVQA训练和验证集的固有偏置,并对测试集问题进行人工重写,以提供多样化和自然的问题形式。数据集的问题数量从9129增加到211572,词汇量也从93增加到465,更好地反映现实世界的语言变异性。FortisAVQA旨在诊断和改进音频视觉问答模型的鲁棒性,特别是在处理分布内和分布外样本时的表现。
FortisAVQA is a dataset specifically designed to evaluate the robustness of audio-visual question answering (AVQA) models. It is constructed through two stages: first, rewriting the test split questions from the publicly released MUSIC-AVQA dataset to enhance diversity; second, introducing question distribution shifts based on answer distributions to enable fine-grained robustness evaluation. This dataset retains the inherent biases of the training and validation splits of the original MUSIC-AVQA dataset, and manually rewrites the test split questions to provide diverse and natural question formats. The number of questions in the dataset has increased from 9129 to 211572, and the vocabulary size has also grown from 93 to 465, which better reflects real-world linguistic variability. FortisAVQA aims to diagnose and improve the robustness of audio-visual question answering models, particularly their performance when handling in-distribution and out-of-distribution samples.

- 1FortisAVQA and MAVEN: a Benchmark Dataset and Debiasing Framework for Robust Multimodal Reasoning西安交通大学 · 2025年



