遇见数据集

Sentinel-2 Surface Water Segmentation Dataset for the Southern Iraqi Marshes (2021)

收藏
Zenodo2026-02-10 更新2026-05-26 收录
官方服务:

资源简介:

Dataset Description This dataset contains Sentinel-2 satellite imagery of the Southern Iraqi Marshes, prepared for training and evaluating fully convolutional neural networks (FCNs) for binary semantic segmentation of surface water. The dataset is intended for research and educational use in remote sensing, hydrology, and machine-learning-based water body mapping, particularly for benchmarking deep learning models on medium-resolution multispectral imagery. Data Sources Sentinel-2 Level-2A surface reflectance imagery (Copernicus / ESA), accessed and exported using Google Earth Engine JRC Global Surface Water (GSW) Yearly History dataset for water extent labels Image Preprocessing Sentinel-2 imagery was exported as four 10 m spatial resolution bands: B4 – Red B3 – Green B2 – Blue B8 – Near-Infrared The bands were exported as 8-bit GeoTIFF files and subsequently processed offline. All image–mask pairs were spatially aligned to ensure identical coordinate reference systems, transforms, and pixel grids. Where required, masks were resampled to the image grid using nearest-neighbour reprojection to preserve discrete class labels. Each image–mask pair was tiled into non-overlapping 512 × 512 pixel patches. Image tiles were then min–max normalised on a per-tile basis and stored as floating-point arrays in the range [0, 1]. Final image tiles have shape (512, 512, 4) and dtype float32. Ground-Truth Water Masks Water masks were derived from the JRC Global Surface Water Yearly History product. Exported water layers were binarised such that: Water pixels → 1 Background (non-water) pixels → 0 Masks were saved as NumPy arrays with shape (512, 512, 1) and dtype uint8. This release uses 2021 data only, as configured in the preprocessing pipeline. Dataset Splitting and Balancing To mitigate severe class imbalance commonly present in water segmentation tasks, tiles were categorised based on water coverage fraction: Empty: 0% water Mixed: >0% and <30% water Water-heavy: ≥30% water A balanced subset of tiles was sampled into training, validation, and test splits using fixed proportions of empty, mixed, and water-heavy tiles. Splits are non-overlapping at the tile level, ensuring no spatial leakage between subsets. Validation and test sets were additionally constrained to include a minimum number of water-heavy tiles where available. Final split sizes: Training set: 250 image–mask pairs Validation set: 100 image–mask pairs Test set: 20 image–mask pairs Intended Use This dataset is designed for: Training and evaluation of semantic segmentation models (e.g. U-Net, SegNet, Attention-U-Net) Research on water body detection, wetland monitoring, and environmental change analysis Benchmarking machine learning pipelines on multispectral satellite imagery Licence and Usage This dataset is released for non-commercial research and educational purposes. Users should cite the original data sources (Copernicus Sentinel-2 and JRC Global Surface Water) when using this dataset in academic work.

数据集说明 本数据集包含伊拉克南部沼泽的哨兵-2(Sentinel-2)卫星影像,用于训练和评估全卷积神经网络(FCNs)以完成地表水体的二元语义分割任务。 本数据集旨在服务于遥感、水文学以及基于机器学习的水体制图领域的研究与教学工作,尤其适用于在中分辨率多光谱影像上对深度学习模型进行基准测试。 数据来源 哨兵-2 Level-2A地表反射率影像(哥白尼计划/欧洲空间局(ESA)),通过谷歌地球引擎(Google Earth Engine)获取并导出; JRC全球地表水(GSW)年度历史数据集,用于生成水体范围标签。 影像预处理 哨兵-2影像被导出为4个10米空间分辨率波段: B4——红光波段 B3——绿光波段 B2——蓝光波段 B8——近红外波段 上述波段以8位GeoTIFF格式导出,随后进行离线处理。所有影像-掩码对均进行空间对齐,以确保坐标参考系统、转换参数与像素网格完全一致。必要时,采用最近邻重采样方法将掩码重采样至影像网格,以保留离散类别标签。 每一组影像-掩码对被切割为互不重叠的512×512像素斑块。随后,影像斑块按每个斑块单独进行最小-最大归一化,存储为取值范围[0, 1]的浮点型数组。最终影像斑块的形状为(512, 512, 4),数据类型为float32。 真值水体掩码 水体掩码源自JRC全球地表水年度历史产品。导出的水体图层被二值化处理,规则如下: 水体像素 → 1 背景(非水体)像素 → 0 掩码以NumPy数组格式存储,形状为(512, 512, 1),数据类型为uint8。 本版本仅使用2021年的数据,预处理流程已对此完成配置。 数据集划分与平衡 为缓解水体分割任务中普遍存在的严重类别不平衡问题,斑块按照水体覆盖占比进行分类: 无水体:水体占比0% 混合类别:水体占比>0%且<30% 富水体类别:水体占比≥30% 采用固定的无水体、混合类别与富水体斑块比例,从斑块中采样得到平衡子集,划分为训练集、验证集与测试集。各子集在斑块级别上互不重叠,确保子集之间不存在空间信息泄露。验证集与测试集还额外设置约束,在数据可用的情况下保证包含至少一定数量的富水体斑块。 最终划分规模如下: 训练集:250组影像-掩码对 验证集:100组影像-掩码对 测试集:20组影像-掩码对 预期用途 本数据集旨在用于以下场景: 1. 语义分割模型(如U-Net、SegNet、注意力U-Net(Attention-U-Net))的训练与评估 2. 水体检测、湿地监测与环境变化分析相关研究 3. 在多光谱卫星影像上对机器学习流水线进行基准测试 许可与使用规范 本数据集仅面向非商业性研究与教学用途发布。若在学术工作中使用本数据集,用户需引用原始数据来源(哥白尼计划哨兵-2与JRC全球地表水数据集)。

提供机构:
Zenodo
创建时间:
2026-02-10
二维码
社区交流群
二维码
科研交流群
商业服务