基于街景图像的合成视觉问答数据集
收藏资源简介:
该数据集由韩国科学技术院的研究人员构建,包含来自五个全球城市的50,000张街景图像。通过对这些图像进行语义分割、深度估计和目标检测,研究人员提取了场景属性,并基于这些属性生成了大量问答对。数据集分为感知问答和组合问答两种类型,并进一步将问答对的答案转换为“思维链”形式,以便更好地评估模型的推理过程。该数据集旨在帮助视觉-语言模型更好地理解和解释城市街景,并推动相关领域的研究和应用。
This dataset was constructed by researchers from the Korea Advanced Institute of Science and Technology (KAIST). It comprises 50,000 street-view images collected from five global cities. Using semantic segmentation, depth estimation and object detection on these images, the researchers extracted scene attributes and generated a large number of question-answer (QA) pairs based on these attributes. The dataset is divided into two categories of QA pairs: perceptual QA and compositional QA. Furthermore, the answers to these QA pairs are converted into the "chain-of-thought (CoT)" format to facilitate better evaluation of model reasoning processes. This dataset is designed to help vision-language models better understand and interpret urban street scenes, and to advance research and applications in relevant fields.




