SuperRS-VQA, HighRS-VQA
收藏资源简介:
GeoLLaVA-8K数据集是一个专注于超高清遥感场景的多模态大型语言模型,能够处理高达8K分辨率的输入。该数据集基于SuperRS-VQA和HighRS-VQA构建,包含22个真实世界的子任务,是目前为止图像尺寸最大的遥感视觉语言数据集。数据集的创建过程包括半自动化的标注流程和基于影响力的数据选择方法,旨在解决超高清遥感图像中图像-文本训练数据的稀缺问题。该数据集的应用领域是遥感数据处理,旨在解决超高清遥感任务中的性能限制问题。
GeoLLaVA-8K dataset is a multimodal large language model focused on ultra-high-definition remote sensing scenarios, capable of processing inputs with resolutions up to 8K. Built upon SuperRS-VQA and HighRS-VQA, this dataset includes 22 real-world subtasks, and it is the largest remote sensing vision-language dataset in terms of image size to date. The creation process of this dataset involves a semi-automated annotation workflow and an influence-based data selection method, aiming to solve the scarcity of image-text training data for ultra-high-definition remote sensing images. The application field of this dataset is remote sensing data processing, with the purpose of addressing performance bottlenecks in ultra-high-definition remote sensing tasks.
GeoLLaVA-8K数据集概述
数据集基本信息
- 名称:GeoLLaVA-8K
- 类型:超高分率遥感视觉语言数据集
- 分辨率:
- SuperRS-VQA:平均8,376×8,378
- HighRS-VQA:平均2,000×1,912
- 数据量:81,367个超高分率图像-文本对
数据集构成
- 数据来源:
- 专家和众包人员手动标注的12K超高分率样本
- 通过GPT-4o半自动生成的100K中高分率(2K×2K)样本
- 数据处理:
- 采用LESS框架进行基于影响力的选择
- 对现有遥感数据集进行去重处理
关键特性
- 遥感图像低语义密度问题:
- 背景标记占比高达73.14%
- 目标标记仅占26.5%,但对性能影响显著
- 创新方法:
- 背景标记剪枝
- 锚定标记选择
相关资源
- 论文:arXiv:2505.21375
- 模型:HuggingFace模型
- 数据集:HuggingFace数据集
引用格式
latex @article{wang2025geollava8kscalingremotesensingmultimodal, title={GeoLLaVA-8K: Scaling Remote-Sensing Multimodal Large Language Models to 8K Resolution}, author={Fengxiang Wang and Mingshuo Chen and Yueying Li and Di Wang and Haotian Wang and Zonghao Guo and Zefan Wang and Boqi Shan and Long Lan and Yulin Wang and Hongzhen Wang and Wenjing Yang and Bo Du and Jing Zhang}, journal={arXiv preprint arXiv:2505.21375}, year={2025}, }




