CloudSEN12 - a global dataset for semantic understanding of cloud and cloud shadow in Sentinel-2
收藏资源简介:
<strong>Description</strong> CloudSEN12 is a large dataset for cloud semantic understanding that consists of 9880 regions of interest (ROIs). Each ROI has five 5090x5090 meters image patches (IPs) collected on different dates; we manually choose the images to guarantee that each IP inside an ROI matches one of the following cloud cover groups: - clear (0%) - low-cloudy (1% - 25%) - almost clear (25% - 45%) - mid-cloudy (45% - 65%) - cloudy (65% >) An IP is the core unit in CloudSEN12. Each IP contains data from Sentinel-2 optical levels 1C and 2A, Sentinel-1 Synthetic Aperture Radar (SAR), digital elevation model, surface water occurrence, land cover classes, and cloud mask results from eight cutting-edge cloud detection algorithms. Besides, in order to support standard, weakly, and self-/semi-supervised learning procedures, cloudSEN12 includes three distinct forms of hand-crafted labelling data: high-quality, scribble, and no annotation. Consequently, each ROI is randomly assigned to a different annotation group: 2000 ROIs with pixel-level annotation, where the average annotation time is 150 minutes (high-quality group). 2000 ROIs with scribble level annotation, where the annotation time is 15 minutes (scribble group). 5880 ROIs with annotation only in the cloud-free (0\%) image (no annotation group). For high-quality labels, we use the Intelligence foR Image Segmentation\cite{iris2019} (IRIS) active learning technology, a system that combines human photo-interpretation and machine learning. For scribble, ground truth pixels were drawn using IRIS but without ML support. Finally, the no annotation dataset is generated automatically, with manual annotation only in the clear image patch. The dataset is already available here: <strong>https://shorturl.at/cgjtz</strong>. Check out our website <strong>https://cloudsen12.github.io/</strong> for examples of how to download the dataset via STAC.
**数据集说明**:CloudSEN12是一款面向云语义理解的大型数据集,共包含9880个感兴趣区域(Regions of Interest, ROIs)。每个感兴趣区域包含5张采集于不同日期的5090×5090米图像块(Image Patches, IPs);我们通过人工筛选图像,确保每个感兴趣区域内的图像块均符合以下云量分组之一: - 无云(0%) - 少云(1%~25%) - 近无云(25%~45%) - 中云量(45%~65%) - 多云(≥65%) 图像块是CloudSEN12的核心单元。每张图像块包含哨兵二号(Sentinel-2)光学级1C与2A数据、哨兵一号(Sentinel-1)合成孔径雷达(Synthetic Aperture Radar, SAR)数据、数字高程模型、地表水发生数据、土地覆盖类别数据,以及8种前沿云检测算法生成的云掩膜结果。 此外,为支持标准监督、弱监督以及自/半监督学习流程,CloudSEN12提供三种不同形式的人工标注数据:高质量标注、涂鸦标注与无标注。因此,每个感兴趣区域会被随机分配至不同的标注组: 1. 2000个带有像素级标注的感兴趣区域(高质量标注组),单样本平均标注时长为150分钟; 2. 2000个带有涂鸦级标注的感兴趣区域(涂鸦标注组),单样本标注时长为15分钟; 3. 5880个仅在无云(0%)图像块上完成标注的感兴趣区域(无标注组)。 对于高质量标注,我们采用了面向图像分割的智能系统(Intelligence for Image Segmentation, IRIS)主动学习技术,该系统结合了人工图像解译与机器学习。涂鸦标注则通过IRIS完成,但未借助机器学习辅助。最后,无标注数据集为自动生成,仅在无云图像块上进行人工标注。 本数据集已在以下链接开放获取:**https://shorturl.at/cgjtz**。如需了解如何通过时空资产目录(SpatioTemporal Asset Catalog, STAC)下载数据集,请访问我们的官网**https://cloudsen12.github.io/**查看示例。



