遇见数据集

Stage 1 MultiTask Training tifs for Multi Modal DL Fusion - Preprocessed 32x 32 native resolution GEDI L2A and L2B Nan fill, padding, min max normalization

收藏
Zenodo2025-12-18 更新2026-05-26 收录
官方服务:

资源简介:

Earth Observation (EO) data were processed in Google Earth Engine. Sentinel-1 imagery underwent pseudo-terrain correction, while optical datasets were cloud-filtered and used for spectral band compositing and index-based feature engineering. For each GEDI L2A and L2B footprint centroid, 32 × 32 pixel GeoTIFF image patches were generated at the native spatial resolution of each EO product. To mitigate spatial autocorrelation, a 1 km grid–based sampling strategy was applied to partition the data into training, testing, and validation subsets (70%, 20%, and 10%, respectively). Image patches were padded where necessary, NaN values were filled, and all bands were min–max normalized. GeoTIFF files follow the naming convention {cov_name}_{split}_{product_field}_{shot_number}_{row_ID}_block_{block_id}.tif (e.g., tcap_train_GEDI_L2B_2021135083748_2002_block_6043889.tif). EO covariates include Sentinel-1 (S1), Sentinel-2 (S2), Sentinel-2 spectral indices (S2IDX), Landsat, Landsat-derived indices, Harmonized Landsat–Sentinel (HLS), Tasseled Cap (TCAP), and a digital elevation model (DEM). Native spatial resolutions are 10 m for S1, 20 m for S2 and S2IDX, and 30 m for Landsat, Landsat indices, HLS, TCAP, and DEM. The number of bands per covariate is as follows: S1 (4), S2 (6), S2IDX (8), Landsat (6), Landsat indices (7), HLS (7), TCAP (6), and DEM (2). The dataset is distributed as two primary archives: GEDI_L2A_processed_32x32_normalized.zip and GEDI_L2B_processed_32x32_normalized.zip. Each archive contains 5,758 normalized GeoTIFF patches with associated product_field metadata. For both GEDI L2A and L2B products, the data are split into 4,280 training samples, 800 testing samples, and 678 validation samples. Passive optical features include median composite spectral bands (Red, Green, Blue, NIR, SWIR1, SWIR2) and derived indices: NDVI ((NIR − Red)/(NIR + Red)), NDWI ((Green − NIR)/(Green + NIR)), MNDWI ((Green − SWIR1)/(Green + SWIR1)), SAVI (1.5 × (NIR − Red)/(NIR + Red + 0.5)), NDMI ((NIR − SWIR1)/(NIR + SWIR1)), NDBI ((SWIR1 − NIR)/(SWIR1 + NIR)), NBR ((NIR − SWIR)/(NIR + SWIR)), and EVI (2.5 × (NIR − Red)/(NIR + 6×Red − 7.5×Blue + 1)). Data sources include Landsat, Sentinel-2, and Harmonized Landsat–Sentinel. Tasseled Cap features derived from Landsat Top-of-Atmosphere data include Brightness, Greenness, Wetness, and higher-order components (4–6), computed as weighted linear combinations of OLI spectral bands. Active microwave features from Sentinel-1 GRD include VV and VH backscatter (dB), the VV/VH ratio, and a normalized difference ratio ((VV − VH)/(VV + VH)). Topographic variables include elevation (meters above mean sea level) and terrain slope (degrees). In addition to the EO patch archives, the release includes a multitask learning implementation package (Multi_Task_Implementation-20251216T163243Z-3-001) supporting model training and reproducibility. This package contains tabular data and metadata files, including df_valid.csv (file paths and associated GEDI measurements), df_valid_multitask.csv (full multitask dataframe, 10.6 MB), dataloader_metadata.pkl (dataset metadata), dataloader_config.pkl (configuration for experiment reproduction), and sample_batches.pt (example batches for testing). The provided DataLoader configuration defines a total of 1,281 samples, partitioned into 958 training, 135 validation, and 188 testing samples. Model targets include GEDI-derived structural metrics (rh, pai, and fhd_normal), and supported input modalities span all EO covariates: 's1', 's2', 'landsat', 'landsatidx', 'hls', 'tcap', 'dem', 's2idx'

提供机构:
Zenodo
创建时间:
2025-12-18
二维码
社区交流群
二维码
科研交流群
商业服务