trust-tad/dual-validation-multispectral
收藏资源简介:
这是一个专为机器学习设计的、多时相的卫星数据集,主要用于云隙填补和分层模型评估。该数据集基于标准分析就绪数据(ARD)严格筛选,包含内在辐射验证,并通过一个新颖的复合难度指数(DI)进行分层,该指数从空间异质性、时间物候变异性和云持续性三个维度评估样本。数据源为MODIS/Terra地表反射率每日L2G全球500米数据(MOD09GA v6.1),辅以ESA WorldCover 2021数据用于空间异质性映射。研究区域为中欧(经度0°–20°E,纬度40°–60°N),时间窗口为2021年6月14日至7月3日(儒略日165–184)。数据集任务为掩码图像建模/云隙填补。数据规模为一个Zarr数据立方体,维度为20×8×4800×4800,代表20天、8个波段和4800×4800的空间网格;其中前7个波段为地表反射率,第8个波段是自定义的多标签可用性掩码,编码了各种大气和质量条件。数据加载器动态处理此掩码,将立方体裁剪为224x224的时空补丁,使用6个活动反射波段作为模型输入。数据集关联模型为trust-tad/dual-validation-multispectral。数据分割策略包括:训练集通过滑动时间窗口(1天重叠)生成以最大化数据可用性;验证集和测试集通过严格非重叠时间窗口和非重叠空间补丁生成;目标帧选择时间窗口内最清晰的帧(可用性τ≥0.85)并人工掩码作为物理重建目标。样本根据难度指数分为三个层次:低难度(DI<0.50,云覆盖最小,低方差)、中难度(0.50≤DI≤0.64)和高难度(DI>0.64,云持续性高,空间/时间复杂性高)。
This is a Machine Learning-ready, multitemporal satellite dataset designed specifically for cloud gap imputation and stratified model evaluation. It acts as the foundational dataset for the framework introduced in A Dual Validation Framework for Curating Machine Learning-Ready Satellite Datasets: A Scalable Pipeline and Stratified Analysis. The dataset is strictly curated from standard Analysis-Ready Data (ARD) to include intrinsic radiometric validation and is stratified using a novel composite Difficulty Index (DI). This index evaluates samples across three dimensions: spatial heterogeneity, temporal phenological variability, and cloud persistence. Source Sensor: MODIS/Terra Surface Reflectance Daily L2G Global 500m (MOD09GA v6.1); Ancillary Data: ESA WorldCover 2021 (for spatial heterogeneity mapping); Study Area: Central Europe (Longitude 0°–20°E, Latitude 40°–60°N); Time Window: June 14 to July 3, 2021 (Julian days 165–184); Task: Masked Image Modeling / Cloud Gap Imputation; Size: A Zarr data cube with dimensions of 20×8×4800×4800 (representing 20 days, 8 bands, and a 4800×4800 spatial grid). The first seven bands are surface reflectance, while the 8th band is a custom multi-label usability mask that encodes various atmospheric and quality conditions. The on-the-fly dataloader dynamically processes this mask to crop the cube into 224x224 spatiotemporal patches, utilizing 6 active reflectance bands for model input. Associated Models: trust-tad/dual-validation-multispectral. Data Splits & Sampling Strategy: To prevent temporal data leakage during model evaluation, the dataset employs strict sampling rules: Training Set generated using a sliding temporal window with a 1-day overlap to maximize data availability; Validation & Test Sets generated using strict non-overlapping temporal windows and non-overlapping spatial patches; Target Frame selected as the clearest frame within a temporal window (usability τ ≥ 0.85) and artificially masked to serve as the physical reconstruction target. Dataset Structure: Samples are categorized into three difficulty strata based on the Difficulty Index (DI): Low Difficulty (DI < 0.50, Minimal cloud cover, low variance), Medium Difficulty (0.50 ≤ DI ≤ 0.64), High Difficulty (DI > 0.64, High cloud persistence, high spatial/temporal complexity).




