Ground-based Pixel-level Cloud Dataset (GPCD)
收藏资源简介:
Using the PVIFS-02 whole-sky imagers, we collected 500,000 independent cloud images from 2021 to 2023, captured in a southern city and a northern city in China. The cloud images collected in southern China are clear, with obvious cloud edges. In contrast, the cloud images from northern China appear relatively blurred. This difference is attributed to the geographical characteristics of northern China, where regions are frequently affected by sand and dust, leading to a certain degree of image blurring. It brings challenges to cloud detection and classification. In order to train and test the algorithms for pixel-level cloud detection and classification, 714 images which contained various types of clouds were selected, and were manually annotated at the pixel-level after being normalized to a resolution of 1024 $\times$ 1024. The annotated dataset, referred to as the Ground-based Pixel-level Cloud Dataset (GPCD), as shown in Fig. \ref{fig3}. This dataset contains two types of annotation files, one with a binarized cloud-sky segmentation and the other classifying clouds into eight categories at the pixel-level according to cloud genera definitions of the World Meteorological Organization (WMO), cloud approximate appearance and sky conditions in practice. Table \ref{tab1} presents the cloud genera and descriptions for each category in GPCD. To further enhance the robustness and applicability of the GPCD, the dataset was subdivided into two region-specific subsets: GPCD-North and GPCD-South. This subdivision was based on the geographical origin of the images, with GPCD-North comprising data collected from northern China and GPCD-South encompassing data from southern China. The rationale behind creating these subsets is to account for regional atmospheric differences that may influence cloud morphology and behavior. By conducting separate analyses on these subsets, we aim to evaluate the performance and robustness of cloud detection and classification algorithms in the context of regional variations. This approach not only allows for a more nuanced understanding of algorithm performance across diverse climatic conditions but also facilitates the testing of transfer learning capabilities.
使用PVIFS-02型全天相机(PVIFS-02 whole-sky imagers),我们于2021年至2023年间,从中国南方与北方两座城市采集了50万张独立云图。采集自中国南方的云图画面清晰,云体边缘轮廓分明;与之相对,北方地区采集的云图则相对模糊。这一差异源于中国北方的地理环境特征:该区域频繁受沙尘天气影响,导致图像出现一定程度的模糊,这为云检测与分类任务带来了挑战。 为训练和测试像素级云检测与分类算法,我们筛选出714张包含各类云系的图像,并将其归一化至1024×1024分辨率后,进行了像素级人工标注。该标注数据集被命名为地基像素级云数据集(Ground-based Pixel-level Cloud Dataset, GPCD),如图 ef{fig3}所示。本数据集包含两类标注文件:一类为云-天空二值分割标注,另一类则依据世界气象组织(World Meteorological Organization, WMO)的云属定义、实际云体外观与天空状况,对像素级云图进行了八分类标注。表 ef{tab1}列出了GPCD中各分类对应的云属与描述信息。 为进一步提升GPCD的鲁棒性与适用性,本数据集依据图像的采集地域,被划分为两个区域专属子集:GPCD-North与GPCD-South。其中GPCD-North包含中国北方采集的图像数据,GPCD-South则涵盖中国南方的采集数据。划分这两个子集的初衷在于,考虑到不同区域的大气差异可能会影响云的形态与特性。通过对这两个子集分别开展分析,我们旨在评估不同区域环境下云检测与分类算法的性能与鲁棒性。该方案不仅能够更细致地探究算法在多样气候条件下的表现,还有助于对迁移学习能力进行测试验证。




