The Kurdistan Burned Area Dataset (KBAD): A Sentinel-2 Deep Learning Benchmark for Multi-Temporal Burn Scar Segmentation in Semi-Arid Mountainous Ecosystems
收藏资源简介:
The Kurdistan Burned Area Dataset (KBAD): A Sentinel-2 Deep Learning Benchmark for Multi-Temporal Burn Scar Segmentation in Semi-Arid Mountainous Ecosystems Overview The Kurdistan Burned Area Dataset (KBAD) is a highly curated, multi-temporal deep learning benchmark dataset engineered to advance automated burn scar semantic segmentation (such as U-Net, Attention U-Net, and Vision Transformer architectures). The dataset focuses on the topographically complex, semi-arid Zagros mountain eco-region within the Kurdistan Region of Iraq (KRI). It provides 317 rigorously quality-controlled satellite scene pairs capturing major wildfire complexes across a ten-year temporal domain (2017–2026). Dataset Contents & Structure The repository contains standardized spatial tensors and tabular metadata ready for machine learning pipelines: Multi-Band Spectral Tensors: Standardized 10-meter resolution grids containing core Sentinel-2 surface reflectance bands (including Visible RGB, Near-Infrared, and Shortwave Infrared channels B11 and B12). Spectral bands are natively co-registered and resampled using bilinear interpolation to preserve radiometric integrity. Topographic & Mask Overlays: Co-registered digital elevation models (30m SRTM DEM) and dynamic categorical Scene Classification Layers (SCL), resampled via nearest-neighbour interpolation to preserve discrete integer codes. Ground-Truth Target Masks: High-fidelity, binary burn scar labels generated via scene-specific adaptive dNBR thresholding (using a 0.20 baseline) and manually cross-verified through meticulous expert photo-interpretation against bitemporal Sentinel-2 true-color and false-color RGB composites. metadata_master.csv: A master matrix detailing unique scene IDs, exact ignition/post-fire acquisition dates, coordinates, regional governorates, and quality metrics. ⚠️ Critical Usage Guideline Users are strongly encouraged to filter the dataset on the master metadata column qa_pass == True before partitioning data into training, validation, or testing splits. The qa_pass flag represents the cumulative outcome of rigorous data engineering filters, isolating pristine, cloud-free, and shadow-corrected scene pairs optimal for deep learning stability. Python import pandas as pd df = pd.read_csv("metadata_master.csv") # Filter for pristine training scenes clean_dataset = df[df['qa_pass'] == True] 📜 Terms of Use & Licensing This dataset is officially published under the Creative Commons Attribution 4.0 International (CC-BY 4.0) license. Permitted Reuse: You are free to share, copy, redistribute, remix, transform, and build upon this material for any purpose, including commercial applications. Attribution Requirement: You must give appropriate credit, provide a link to the license, and indicate if changes were made. You must do so in a reasonable manner, but not in any way that suggests the licensor endorses you or your use. Any academic publication, technical white paper, or commercial product utilizing this dataset must formally cite the associated data descriptor article and link to this Zenodo repository DOI. 📊 Suggested Citation If you utilize this benchmark dataset, its tensor design, or the associated metadata matrix in your research, please cite both this repository DOI and the official peer-reviewed data article.



