RSICap
收藏资源简介:
RSICap是一个专为遥感领域设计的高质量图像描述数据集,由阿里巴巴达摩院创建。该数据集包含2585个人工标注的图像描述,每个描述都提供了丰富的场景和物体信息,如颜色、形状、位置和数量等。RSICap不仅描述了图像的主要场景,还详细描述了物体之间的关系和视觉推理知识。此外,数据集还包括一个评估集RSIEval,用于评估视觉语言模型在遥感图像描述和视觉问答任务上的表现。RSICap和RSIEval的创建旨在推动遥感领域大型视觉语言模型的发展,解决现有数据集在规模和质量上的不足。
RSICap is a high-quality image captioning dataset specifically designed for the remote sensing field, created by Alibaba DAMO Academy. This dataset contains 2585 manually annotated image captions, each providing rich scene and object information such as color, shape, position, quantity and other relevant details. RSICap not only describes the main scenes of the images but also elaborates on the relationships between objects and visual reasoning knowledge. Additionally, the dataset includes an evaluation set named RSIEval, which is used to assess the performance of vision-language models on remote sensing image captioning and visual question answering tasks. The development of RSICap and RSIEval aims to promote the advancement of large-scale vision-language models in the remote sensing domain, addressing the shortcomings of existing datasets in terms of scale and quality.

- 1RSGPT: A Remote Sensing Vision Language Model and Benchmark阿里巴巴达摩院 · 2023年



