SpaRE
收藏资源简介:
SpaRE数据集是由滑铁卢大学Vector Institute的研究人员创建的,旨在增强视觉-语言模型的空间推理能力。数据集包含455,000个样本,包含3.4百万个问答对,由超详细图像描述生成。该数据集利用了DOCCI、PixMo-Cap和Localized Narratives等超详细图像描述数据集,通过LLM技术提取出与空间推理相关的问答对。数据集的创建过程涉及对超详细描述的过滤、提示构建、问答对生成以及质量保证等步骤。SpaRE数据集的应用领域包括机器人、自动驾驶、导航等领域,旨在解决视觉-语言模型在空间推理方面的不足,提高其在实际任务中的表现。
The SpaRE dataset was developed by researchers from the Vector Institute at the University of Waterloo, with the goal of enhancing the spatial reasoning capabilities of vision-language models. It contains 455,000 samples and a total of 3.4 million question-answer pairs generated from ultra-detailed image captions. This dataset leverages existing ultra-detailed image caption datasets including DOCCI, PixMo-Cap, and Localized Narratives, and extracts question-answer pairs related to spatial reasoning using large language model (LLM) technologies. The dataset creation workflow includes steps such as filtering ultra-detailed captions, prompt construction, question-answer pair generation, and quality assurance. The application fields of the SpaRE dataset cover robotics, autonomous driving, navigation and other domains, aiming to mitigate the shortcomings of vision-language models in spatial reasoning and improve their performance in practical tasks.




