mcosarinsky/NSD-VQA
收藏资源简介:
NSD-VQA是一个大规模视觉问答基准数据集,旨在研究从人类对自然图像的fMRI(功能性磁共振成像)响应中可解码的视觉和语义信息。该数据集基于Natural Scenes Dataset (NSD),并提供了自动生成的问题-答案注释,这些注释与NSD图像相关联。数据集包含约73,000张NSD图像,每张图像平均有约20个问题-答案对,覆盖20个受控语义问题类别,如对象识别、计数、颜色、动作、空间位置、场景理解、人-物交互以及动物、车辆、食物、家电和家居物品等语义类别。该基准设计用于从脑活动中进行细粒度的视觉和语义解码评估,超越单一聚合的VQA分数。注意:此存储库仅包含生成的注释和元数据,不重新分发NSD图像或fMRI记录,用户需根据原始条款单独获取原始NSD数据集。
NSD-VQA is a large-scale visual question answering benchmark for studying what visual and semantic information can be decoded from human fMRI responses to natural images. It is built from the Natural Scenes Dataset (NSD) and provides automatically generated question-answer annotations grounded in NSD images. The dataset contains approximately 73K NSD images, ~20 question-answer pairs per image, and 20 controlled semantic question categories, including object recognition, counting, color, actions, spatial position, scene understanding, human-object interactions, and semantic categories such as animals, vehicles, food, appliances, and household objects. The benchmark is designed for fine-grained evaluation of visual and semantic decoding from brain activity, beyond a single aggregate VQA score. Note: This repository contains only generated annotations and metadata and does not redistribute NSD images or fMRI recordings; users must obtain the original NSD dataset separately under its original terms.



