RS-GPT4V
收藏资源简介:
RS-GPT4V是由中南大学开发的统一多模态指令遵循数据集,专为遥感图像理解设计。该数据集结合了GPT-4V和现有数据集,通过问题-回答对形式统一了任务,如描述、定位等。数据集旨在训练模型理解复杂场景和进行高级视觉推理,包含91,937个训练图像和991,206个问题-回答对,以及15,999个测试图像和258,419个问题-回答对。RS-GPT4V的应用领域广泛,包括图像描述、视觉问答和复杂场景理解,旨在解决遥感图像理解中的复杂性和多样性问题。
RS-GPT4V is a unified multimodal instruction-following dataset developed by Central South University, specifically designed for remote sensing image understanding. This dataset integrates GPT-4V and existing datasets, unifying tasks including image description, object localization and others in the form of question-answer pairs. It aims to train models to comprehend complex scenes and perform advanced visual reasoning, containing 91,937 training images, 991,206 training question-answer pairs, 15,999 test images and 258,419 test question-answer pairs. RS-GPT4V has a wide range of application scenarios, including image captioning, visual question answering (VQA) and complex scene understanding, and is intended to address the complexity and diversity issues in remote sensing image understanding.
RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding
数据集概述
RS-GPT4V 是一个集成了视觉和语言数据的高级任务数据集。该数据集通过多模态指令跟随格式,促进对遥感图像的复杂推理和详细理解。
遥感任务和数据的演变
从简单的遥感任务演变为使用多模态数据进行复杂指令任务。
RS-GPT4V 数据集的设计原则和特点
展示了数据集的设计原则,重点关注统一性、多样性、正确性、复杂性、丰富性和鲁棒性。
RS-GPT4V 数据集构建的原则驱动流程
构建过程遵循结构化方法,包括数据收集、指令-响应生成和指令-标注适应。
引用
如果您发现 RS-GPT4V 对您的研究和应用有用,请使用以下 BibTeX 引用:
@ARTICLE{10197260, author={Xu, Linrui and Guo, Wang and Li, Qiujun and Long, Kewang and Zou, Kaiqi and Wang, Yuhan and Li, Haifeng}, title={RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding}, year={2024}, volume={}, number={}, pages={1-14}, journal={arXiv}, doi={https://arxiv.org/abs/2406.12479} }

- 1RS-GPT4V: A Unified Multimodal Instruction-Following Dataset for Remote Sensing Image Understanding中南大学 · 2024年



