xiang709/VRSBench
收藏资源简介:
VRSBench是一个用于遥感图像理解的多功能视觉-语言基准数据集。它包含29,614张遥感图像,每张图像都有详细的描述,52,472个对象引用,以及3,123,221个视觉问答对。这些数据支持广泛的遥感图像理解任务的训练和评估。数据集通过多个数据工程步骤构建,包括属性提取、提示工程、GPT-4推理和人工验证。此外,数据集还支持模型训练,展示了LVMs在遥感图像理解中的潜力,并讨论了数据集的社会影响、偏见、已知限制和未来工作。
VRSBench is a Versatile Vision-Language Benchmark for Remote Sensing Image Understanding. It consists of 29,614 remote sensing images with detailed captions, 52,472 object refers, and 3,123,221 visual question-answer pairs. It facilitates the training and evaluation of vision-language models across a broad spectrum of remote sensing image understanding tasks. The dataset is constructed through multiple data engineering steps, including attribute extraction, prompting engineering, GPT-4 inference, and human verification. Additionally, it supports model training, demonstrating the potential of LVMs in remote sensing image understanding, and discusses the datasets social impact, biases, known limitations, and future work.
VRSBench 数据集概述
数据集基本信息
- 许可证: Creative Commons Attribution Non Commercial 4.0
- 任务类别: 视觉问答、文本生成
- 语言: 英语
- 名称: VRSBench
- 大小类别: 10K<n<100K
- 标签: 遥感、视觉语言模型
数据集内容
- 图像数量: 29,614 张遥感图像
- 对象标注: 52,472 个对象标注
- 视觉问答对: 312,221 对
数据集构建
- 属性提取: 从现有对象检测数据集中提取图像和对象信息。
- 提示工程: 设计指令以提示 GPT-4V 生成详细的图像标题、对象引用和问答对。
- GPT-4 推理: 使用 OpenAI API 自动生成图像标题、对象引用和问答对。
- 人工验证: 通过人工标注者验证 GPT-4V 生成的每个标注。
模型训练
- 基准模型: LLaVA-1.5, MiniGPT-v2, Mini-Gemini, GeoChat
- 微调: 在 RSVBench 数据集上对每个模型进行 5 个周期的微调,使用 LoRA 微调,秩为 64。
数据集影响
- 社会影响: 支持高级视觉语言模型的训练和评估,提升其在遥感中的应用能力。
- 偏见讨论: 尽管通过人工验证确保高质量标注,但视觉数据的解释可能存在主观偏见。
- 其他已知限制: 地理多样性受限于 DOTA-v2 和 DIOR 数据集覆盖的区域。
许可证信息
- 许可证: Creative Commons Attribution Non Commercial 4.0
未来工作
- 扩展计划: 计划将 VRSBench 扩展到包括红外图像、多光谱和超光谱图像、合成孔径雷达(SAR)图像和时间数据集在内的多种遥感数据类型。
引用信息
bibtex @misc{li2024vrsbench, title={VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding}, author={Xiang Li, Jian Ding, Mohamed Elhoseiny}, year={2024}, eprint={xxx}, archivePrefix={arXiv}, primaryClass={cs.CV} }




