VRSBench
收藏资源简介:
VRSBench是由阿卜杜拉国王科技大学创建的一个大型多任务遥感图像理解数据集,包含29,614张图像,每张图像配有详细的人工验证描述、52,472个对象引用和123,221个问答对。数据集通过半自动数据收集流程创建,包括属性提取、提示工程、GPT-4推理和人工验证,确保高质量和大规模。VRSBench旨在推动遥感领域中视觉-语言模型的发展,支持图像描述、视觉定位和视觉问答等多种任务,解决现有数据集在详细对象信息、质量控制和多任务适应性方面的不足。
VRSBench is a large-scale multi-task remote sensing image understanding dataset developed by King Abdullah University of Science and Technology (KAUST). It comprises 29,614 images, each paired with detailed human-validated descriptions, 52,472 object references, and 123,221 question-answer pairs. The dataset is constructed via a semi-automated data collection pipeline that includes attribute extraction, prompt engineering, GPT-4 inference, and human validation, ensuring both high quality and large-scale coverage. VRSBench aims to advance the development of vision-language models in the remote sensing domain, supports multiple tasks such as image captioning, visual grounding, and visual question answering, and addresses the limitations of existing datasets in terms of detailed object information, quality control, and multi-task adaptability.

- 1VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding阿卜杜拉国王科技大学 · 2024年



