遇见数据集

VQA-MHUG

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

我们提出了 VQA-MHUG - 一个新颖的 49 人数据集,包含使用高速眼动追踪器收集的视觉问答 (VQA) 期间的图像和问题的多模态人类注视。我们使用我们的数据集来分析由五个最先进的 VQA 模型学习的人类和神经注意力策略之间的相似性:具有网格或区域特征的调制共同注意力网络 (MCAN)、Pythia、双线性注意力网络 (BAN) ,以及多模态分解双线性池网络 (MFB)。虽然之前的工作集中在研究图像模态,但我们的分析首次表明,对于所有模型,与人类对文本的注意力的更高相关性是 VQA 性能的重要预测指标。这一发现指出了提高 VQA 性能的潜力,同时要求进一步研究神经文本注意机制及其与视觉和语言任务架构的集成,包括但可能超出 VQA。

We introduce VQA-MHUG, a novel 49-participant multi-modal human gaze dataset that captures both image and question modalities during visual question answering (VQA) tasks, collected using high-speed eye trackers. We employ this dataset to analyze the similarity between human and neural attention strategies learned by five state-of-the-art VQA models: Modulated Co-Attention Network (MCAN) with grid or regional features, Pythia, Bilinear Attention Network (BAN), and Multimodal Factorized Bilinear Pooling network (MFB). While prior work has focused exclusively on image modality, our analysis is the first to demonstrate that higher correlation with human attention to text serves as a critical predictor of VQA performance across all evaluated models. This finding highlights the potential for improving VQA performance, while also calling for further research into neural textual attention mechanisms and their integration with visual and language task architectures, including but not necessarily limited to VQA.

提供机构:
OpenDataLab
创建时间:
2022-09-01
搜集汇总
数据集介绍
VQA-MHUG 数据集图片
背景与挑战
背景概述
VQA-MHUG是一个包含49人多模态人类注视数据的视觉问答数据集,用于分析五种先进VQA模型与人类注意力策略的相似性。研究发现,文本注意力与人类注意力的更高相关性是预测VQA性能的重要指标,这为提高模型性能提供了新方向。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务