遇见数据集

Synchronized Cloud Classification Dataset from FY4A Multispectral Observations with Himawari-8 Labels

收藏
Zenodo2025-09-24 更新2026-05-26 收录
官方服务:

资源简介:

SynCloud-FY4A is a benchmark dataset for cloud classification constructed from FY-4A AGRI multispectral observations with labels derived from Himawari-8 L2 cloud classification products. Covering the region from 5°N–45°N and 90°E–130°E over 2018–2023, the dataset includes 12,000 collocated pairs of FY-4A inputs and Himawari-8 labels. After preprocessing (radiometric calibration, parallax and geometric correction, projection matching), the data are provided at a 0.05°×0.05° grid resolution, resulting in images of size 800×800 (40°×40°). Users can further process these data according to their research needs. In our cloud classification experiments, the original images were cropped into patches of 224×224 for network training. The Himawari-8 products distinguish nine cloud types (Ci, Cs, Dc, Ac, As, Ns, Cu, Sc, St), serving as reliable annotations. This dataset can be widely used for satellite-based cloud classification, meteorological forecasting, and climate studies. Original Data Sources: FY4A AGRI L1: https://satellite.nsmc.org.cn/dataportal/cn/home/index.html Example: FY4A-_AGRI--_N_DISK_1047E_L1-_FDI-_MULT_NOM_20190101000000_20190101001459_4000M_V0001.HDF Himawari-8 L2: https://www.eorc.jaxa.jp/ptree/ Example: NC_H08_20190101_0000_L2CLP010_FLDK.02401_02401.nc Dataset Structure: data/ |-- images/ |-- training/ |-- test/ |-- annotations/ |-- training/ |-- test/ Data Processing: Preprocessing: Radiometric calibration, parallax correction, geometric correction, projection transformation, and spatiotemporal matching. Dataset Split: Randomly divided into training and test sets (ratio 9:1). Model Training: Used for training the cloud classification network proposed in the associated study, as well as baseline comparison networks. Model Testing & Evaluation: Performance validated using the test set. Usage Notes: After extracting the archive, files starting with himawari-8 contain cloud-type labels, and files starting with fy4a contain multispectral observations. Users should organize observation and label data into their respective directories and run the dataset split script. Training data should be saved under ../images/training/ and labels under ../annotations/training/; similarly for the test set. Currently, only 2019 FY4A observations and Himawari-8 labels are included. Data from other years will be uploaded gradually and linked to this dataset.

提供机构:
Zenodo
创建时间:
2025-09-24
二维码
社区交流群
二维码
科研交流群
商业服务