Globe230k: A Benchmark Dense-Pixel Annotation Dataset for Global Land Cover Mapping
收藏资源简介:
We (Intelligent Mining and Analysis of Remote Sensing big data, IMARS) create a large-scale annotated dataset (Globe230k) for land use/land cover (LULC) mapping, which is annotated on Google Earth image of 1 m spatial resolution. Globe230k is annotated by numerous experts and students major in survey and mapping after necessary training, through visual interpretation on very high-resolution images, as well as in-situ field survey, under the guidance of the organized annotation pipeline. Globe230k has three superiorities: 1) Large scale: the Globe230k includes 232,819 annotated images with the size of 512x512 and spatial resolution of 1 m, with more than 3x1010 annotated pixels, and it includes 10 first-level categories. 2) Rich diversity: the annotated images are sampled from worldwide regions, with coverage area of over 60,000 km2, indicating a high variability and diversity. Besides, in order to ensure the category balance, we intentionally give more chance to the rare categories to be sampled, such as wetland, ice/snow, etc. 3) Multi-modal: Globe230k not only contains RGB bands, but also include other important features for Earth system research, such as Normalized differential vegetation index (NDVI), digital elevation model (DEM), vertical-vertical polarization (VV) bands, vertical-horizontal polarization (VH) bands, which can facilitate the multi-modal data fusion research. Due to the large size of the multi-modal dataset (DEM 1.91G, NDVI 164G, VVVH 372G), these dataset are stored on Baidu Yunpan, the download link is :https://pan.baidu.com/s/12AKbiqOXSf4fnm7mYkCE0g?pwd=230k, the extraction code is 230k. The image patches and their corresponding annotated patches are respectively stored in "image_patch.zip" and "label_patch.zip" file. The RGB image is in forms of ".jpg", with size of 512x512, the pixel value is ranged from 0-255. The annotated patches is in forms of ".png", also with size of 512x512, the pixel value is ranged from 1-10, which respectively represent 1#cropland, 2#forest, 3#grass, 4#shrubland, 5#wetland, 6#water, 7#tundra, 8#impervious, 9#bareland, 10#ice/snow. The corresponding DEM, NDVI and VVVH patches are all in form of ".tif", with size of 512x512 (due to the different resolution of DEM, NDVI and VVVH patches, they are all uniformly resized to the same scale as the image patch). The total 232,819 pairs are officially divided into training set, validation set, and test set, based on ratio of 7:1:2, which can be find in "train_num.txt","val_num.txt","test_num.txt" file. Based on this division, the official baseline accuracy of several state-of-the-art semantic segmentation can be found in the related arcticle (https://spj.science.org/doi/10.34133/remotesensing.0078). We hope it can be used as a benchmark to promote further development of global land cover mapping and semantic segmentation algorithm development.
我们(遥感大数据智能挖掘与分析团队,Intelligent Mining and Analysis of Remote Sensing big data,缩写IMARS)构建了一款用于土地利用/土地覆盖(Land Use/Land Cover, LULC)制图的大规模标注数据集Globe230k,该数据集基于1米空间分辨率的谷歌地球(Google Earth)影像进行标注。Globe230k由众多经过专业培训的测绘类专家与学生,在标准化标注流程的指导下,通过高分辨率影像目视解译结合实地野外勘测完成标注。Globe230k具备三大优势: 1) 大规模性:Globe230k包含232819张512×512分辨率的标注影像,标注像素总量超过3×10¹⁰,涵盖10个一级类别。 2) 丰富多样性:标注影像采样自全球各地,覆盖面积超过60000平方千米,具备极高的类别变异性与多样性。此外,为保障类别平衡,我们特意提升了湿地、冰雪等稀有类别的采样权重。 3) 多模态特性:Globe230k不仅包含RGB波段影像,还涵盖了地球系统研究所需的其他重要特征,例如归一化差分植被指数(Normalized Differential Vegetation Index, NDVI)、数字高程模型(Digital Elevation Model, DEM)、垂直-垂直极化(Vertical-Vertical, VV)波段与垂直-水平极化(Vertical-Horizontal, VH)波段,可助力多模态数据融合研究。由于多模态数据集体量较大(DEM:1.91GB,NDVI:164GB,VVVH:372GB),相关数据存储于百度网盘,下载链接为:https://pan.baidu.com/s/12AKbiqOXSf4fnm7mYkCE0g?pwd=230k,提取码为230k。 图像块与对应的标注块分别存储于"image_patch.zip"与"label_patch.zip"文件中。RGB影像采用.jpg格式,分辨率为512×512,像素值范围为0-255。标注块采用.png格式,分辨率同样为512×512,像素值范围为1-10,分别对应:1#耕地、2#森林、3#草地、4#灌丛、5#湿地、6#水域、7#苔原、8#不透水面、9#裸地、10#冰雪。对应的DEM、NDVI与VVVH影像块均采用.tif格式,分辨率为512×512(考虑到DEM、NDVI与VVVH影像原始分辨率存在差异,所有数据均统一重采样至与图像块相同的尺度)。 全部232819对数据按照7:1:2的比例划分为训练集、验证集与测试集,相关划分信息可在"train_num.txt"、"val_num.txt"、"test_num.txt"文件中查看。基于该划分标准,多款当前领先的语义分割模型的官方基准精度可参见相关研究论文(https://spj.science.org/doi/10.34133/remotesensing.0078)。 我们期望该数据集可作为基准数据集,推动全球土地覆盖制图与语义分割算法的进一步发展。



