SENSERO: Semantic Earth-observation and Natural-language Scene Explanation for Romania (2018)
收藏资源简介:
SENSERO is a multimodal, multi-scale dataset that pairs Sentinel-2 imagery with CORINE Land Cover 2018 (CLC2018) reference semantics over Romania. Each of the 10,000 geographically fixed patch locations is provided at eight co-registered spatial scales with identical centers, and is annotated with five parallel, salience-ordered multi-label sets. The dataset supports multi-label land-cover classification, image–text retrieval, the study of spatial-scale effects on semantic complexity, and — as a downstream use — Sentinel-2 caption generation. What makes SENSERO new Three design choices set SENSERO apart from earlier single-scale, single-taxonomy patch benchmarks: Eight co-registered scales with identical patch centers (64–336 px; 640 × 640 m to 3360 × 3360 m at 10 m GSD). Because the centers are shared, spatial scale becomes an explicit, controllable variable across the same locations rather than a fixed design parameter — a multi-scale axis single-resolution datasets cannot offer. Five parallel label spaces per patch — the three CORINE Land Cover hierarchy levels, a BigEarthNet-compatible scheme, and a new SENSERO scheme — so a single patch can be interpreted at whatever semantic granularity a task needs, from coarse groups to fine classes. Salience-ordered labels — within each set, labels are ranked by prominence via center-weighting, with anthropogenic classes prioritized when present, approximating how a human observer would describe the scene (naming the dominant, central, often human-made feature first) rather than handing over an unordered set. This adds a layer of information absent from conventional multi-label annotations. One label space follows the BigEarthNet nomenclature (Sumbul et al., 2021), so BigEarthNet-compatible code and models can be reused with minor data-loader adjustments (SENSERO uses one multi-band GeoTIFF per patch instead of one file per band, and a different folder layout). SENSERO is not a drop-in replacement for BigEarthNet, however: the eight co-registered scales and the four additional label spaces give SENSERO a substantially broader scope. Dataset composition 10,000 patch locations, each extracted at the eight spatial scales above. Per patch: Multispectral surface reflectance — 16-bit cube with the 12 surface-reflectance bands of the Sentinel-2 L2A product (band order: B1,B2,B3,B4,B5,B6,B7,B8,B8A,B9,B11,B12; the cirrus band B10 is not part of L2A and is therefore absent). True-color composite (B4-B3-B2) and false-color NIR composite (B8-B4-B3). CLC2018 reference raster in numeric (16-bit, Level-3 class codes) and colorized (8-bit, official Copernicus palette) form. Five multi-label sets (defined below, with an illustrative example). Labeling schemes Each patch carries five multi-label sets, designated caption1–caption5. Within a set, labels are comma-separated, lowercase, and free of internal punctuation, and are ordered by salience: a center-weighted strategy analyzes the central ~20% of the patch first and prioritizes anthropogenic classes when present, while classes occupying less than 2% of the patch area are omitted. These label sets are scale-specific — patch centers are shared across scales, but as the extent grows, a location’s labels typically expand and re-order, with dominant central classes staying at the front and newly included peripheral classes appended — so each scale’s metadata file records its own caption1–caption5 values: caption1 — CORINE Land Cover Level 1 (5 classes) caption2 — CORINE Land Cover Level 2 (15 classes) caption3 — CORINE Land Cover Level 3 (44 classes; 34–35 present in SENSERO) caption4 — BigEarthNet-compatible nomenclature (the 19-class scheme of Sumbul et al., 2021); some CLC Level-3 classes are unlabeled under this scheme. caption5 — the custom SENSERO scheme (11 classes; see below; also has some unlabeled classes). Within each metadata file the sets are stored in columns caption1–caption5. They contain ordered multi-label lists rather than free-text sentences, and double as captions for caption-generation tasks. The SENSERO classification scheme — our contribution While the first four schemes reuse established nomenclatures (the three CLC levels and the BigEarthNet 19-class scheme), caption5 is an original, compact land-cover vocabulary that we introduce with this dataset. It re-aggregates the CLC Level-3 classes into 11 intuitive, plain-language classes chosen so that the labels read naturally in scene descriptions, while remaining deterministically derived from, and fully traceable to, CLC2018. The grouping draws on our experience with Sentinel-2 imagery over Romania and follows two principles. First, classes that are spectrally and ecologically alike in this temperate setting are merged: vineyards and fruit-tree/berry plantations become permanent crops; the natural-grassland, heath, sclerophyllous and transitional-woodland classes all become shrubland; and inland and marine waters (likewise inland and maritime wetlands) are each collapsed into a single class. Second, classes that are marginal across Romania — rarely present, or rarely dominant at the patch level (port areas, airports, beaches, bare rocks) — are left unlabeled. The full mapping is: SENSERO class Aggregated CLC Level-3 classes urban fabric continuous urban fabric (111), discontinuous urban fabric (112), industrial or commercial units (121) quarries or dump sites mineral extraction sites (131), dump sites (132), construction sites (133) agriculture non-irrigated arable land (211), permanently irrigated land (212), rice fields (213) permanent crops vineyards (221), fruit trees and berry plantations (222) pastures pastures (231) mixed farmland complex cultivation patterns (242), land principally occupied by agriculture with significant areas of natural vegetation (243) forests broad-leaved forest (311), coniferous forest (312), mixed forest (313) shrubland natural grasslands (321), moors and heathland (322), sclerophyllous vegetation (323), transitional woodland-shrub (324) bare ground sparsely vegetated areas (333) wetlands inland marshes (411), salt marshes (421) water bodies water courses (511), water bodies (512), coastal lagoons (521), sea and ocean (523) unlabeled all other CLC Level-3 classes, including road and rail networks (122), port areas (123), airports (124), green urban areas (141), sport and leisure facilities (142), beaches, dunes, sands (331). Relative to the other schemes, caption5 is one of the coarsest and the most caption-friendly: it collapses all forest types into forests, all scrub/grassland/heath classes into shrubland, and the arable classes into agriculture. It also recovers the mineral-extraction, dump and construction classes that BigEarthNet leaves unlabeled (as quarries or dump sites). Like BigEarthNet, however, it leaves several artificial classes (roads, ports, airports, green urban areas, sport and leisure facilities) and bare rocks unlabeled, and unlike BigEarthNet it does not label beaches/dunes/sands — so for guaranteed full class coverage the CLC-hierarchy sets (caption1–caption3) should be used. Like all other label sets, caption5 is generated deterministically by the rule-based pipeline and carries no model-provider restrictions. Example — the patch S2A_MSIL2A_20180809T090551_N0500_R050_T35TNM_1049_1798 at 336 px scale, under all five schemes: CLC Level-3 codes 112, 211, 222, 231, 242, 243, 311 caption1 — CLC L1 forest and semi-natural zones, agricultural areas, artificial surfaces caption2 — CLC L2 forests, arable land, permanent crops, heterogeneous agricultural areas, pastures, urban fabric caption3 — CLC L3 broad-leaved forest, non-irrigated arable land, fruit trees and berry plantations, land principally occupied by agriculture with significant areas of natural vegetation, pastures, discontinuous urban fabric, complex cultivation patterns caption4 — BigEarthNet broad-leaved forest, arable land, permanent crops, land principally occupied by agriculture with significant areas of natural vegetation, pastures, urban fabric, complex cultivation patterns caption5 — SENSERO forests, agriculture, permanent crops, mixed farmland, pastures, urban fabric This shows the SENSERO mapping in action: the two heterogeneous agriculture classes in caption3 — complex cultivation patterns (242) and land principally occupied by agriculture with significant areas of natural vegetation (243) — collapse into the single mixed farmland label in caption5, so the seven CLC Level-3 codes resolve to six SENSERO labels. Annotation provenance All label sets were generated deterministically from the CLC2018 reference by a rule-based pipeline — no LLM or AI annotation system was used. They are grounded in an authoritative source, free of hallucinated text, and carry no model-provider output restrictions, so the dataset can be used freely to train and evaluate multi-label, retrieval, and vision-language models — including as supervision for caption generation — with fully reproducible provenance. Train / validation / test split A single partition is provided — 70% train / 15% validation / 15% test (7,000 / 1,500 / 1,500 patches) — stored in the split column. The split was built at the canonical 128 px scale and propagated by patch identifier to all eight scales, so subset membership is identical across resolutions. The procedure used tile-level multilabel stratified sampling (so every Sentinel-2 tile contributes proportionally to all subsets), deterministic rare-class placement (classes present in 1–3 patches are forced so that every class appears in every subset), and a spatial buffer of ≥ 5,000 m between validation/test centroids to reduce spatial autocorrelation leakage. Class proportions differ by only 2–3% across scales, and no rare class is absent from any subset at any resolution. Data organization Distributed as 16 ZIP archives (8 scales × 2 formats; total ≈ 93Gb): geotiff_patch_[size].zip — LZW-compressed GeoTIFF, full spectral + geospatial information. png_patch_[size].zip — 8-bit PNG composites for lightweight visualization and benchmarking. Within an archive, data live under [GeoTiff|PNG]/Patch_<size>/ and are grouped by modality, then by Sentinel-2 scene identifier: Multispectral/ — 12-band reflectance cube (GeoTIFF only) TrueColor_RGB/ — B4-B3-B2 composite (suffix _B432) FalseColor_NIR/ — B8-B4-B3 composite (suffix _B843) CLC_reference/ — 16-bit CLC Level-3 code raster (suffix _CLC, GeoTIFF only) CLC_reference_RGB/ — colorized CLC raster (suffix _CLC_RGB) The PNG tree contains only the composite modalities (TrueColor_RGB, FalseColor_NIR, CLC_reference_RGB). File names follow [Satellite]_[ProductLevel]_[AcquisitionTime]_[Baseline]_[Orbit]_[Tile]_[X]_[Y]_[BandCombo].[ext], where X/Y are the patch top-left column/row within the parent MGRS tile; the shared base name (without the band-combo suffix) is the per-patch join key across all modalities and scales. Each Patch_<size>/ directory contains a metadata_<size>.parquet table that links imagery, labels, and the split. The schema is identical across scales: BaseFolder — Sentinel-2 scene identifier BaseFilename — patch base filename (join key across modalities and scales) CLC_codes — comma-separated list of CLC Level-3 codes present in the patch caption1–caption5 — the five salience-ordered multi-label sets split — train / validation / test assignment A companion vector file sensero_patch_footprints.gpkg (CRS: WGS-84 / EPSG:4326) provides max-scale (3360 × 3360 m) patch footprints and serves as a spatial index across all scales (shared centers). Its attribute table carries the patch identifier, the relative path to the canonical multispectral GeoTIFF, the five label sets, the CLC Level-3 codes, and the split assignment. GeoPackage is used instead of Shapefile because the latter's 254-character field limit would truncate the longer label sets. A full technical description is provided in sensero_technical_guide.pdf. Geographic and thematic coverage Romania offers strong territorial diversity — agricultural plains, forested mountains, dense urban areas, and the Danube Delta wetlands. The dataset spans 34 of the 44 CLC2018 Level-3 classes (35 at the largest scale), with patch centers manually placed for uniform national coverage and balanced class representation, while excluding clouds, haze, and acquisition artefacts. Technical details Source: 38 Sentinel-2 L2A scenes, acquired 9 Aug – 14 Oct 2018 (chosen to align with the CLC2018 reference year and to limit cloud cover and phenological variation). Atmospheric correction via Sen2Cor. Reference: CORINE Land Cover 2018, version v2020_u1. Processing: ESA SNAP v12.0.1; reference maps via QGIS 3.40 + GDAL 3.8 (Python). Resampling: all optical bands resampled to 10 m by bicubic interpolation, using B2 as geometric reference. CLC2018 rasterized at 10 m and reprojected to each Sentinel-2 tile's CRS. CRS: rasters in per-scene UTM zone 34N (EPSG:32634) or 35N (EPSG:32635); CLC2018 native EPSG:3035; sensero_patch_footprints.gpkg in WGS-84 (EPSG:4326). Composite stretch: true-color per-band stretch (min 0.01, max 0.50, gamma 0.45); false-color (min 0.01, max 0.60, gamma 0.55); fixed, scene-invariant parameters for cross-scene comparability. Formats: GeoTIFF (LZW, predictor 2), PNG (8-bit), metadata in Apache Parquet. QA: two rounds of visual inspection (clouds/haze/artefacts removed and re-extracted), geometric overlap analysis, automated CLC integrity checks (no-data / incomplete coverage flagged and replaced), and semantic verification on 10% of the dataset, including cross-scale consistency of the salience ordering. At the largest scale, 125 patches have minor partial overlaps (each < 1% of patch area); no overlaps at smaller scales. Usage notes and limitations Captioning as a downstream task. The five label sets are ordered lists of labels, not natural-language sentences. They are directly usable as supervision for caption generation, but a user wanting fluent captions must add a verbalization step. CLC minimum mapping unit (25 ha). Labels inherit CLC2018's 25 ha minimum mapping unit and generalized class boundaries. Combined with the < 2% area threshold, fine-scale features visible in the imagery may not appear in the labels — most relevant at the smallest scales (a 64 px patch is only ~41 ha). Labels are not derived from the imagery. Labels come exclusively from CLC2018 and do not encode visual inference, so context-dependent features absent from CLC are not represented. Unlabeled classes in caption4 and caption5. The BigEarthNet and SENSERO schemes leave some CLC Level-3 classes unlabeled (see above); the CLC-hierarchy sets (caption1–caption3) provide complete coverage. Sensor characteristics. All data inherit Sentinel-2 L2A radiometric/geometric properties; subtle residual differences across scenes may persist despite the standardized pipeline. License and attribution This dataset comprises two layers with distinct licensing. Authors’ contributions. The patch extraction and multi-scale organization, the five label-set (caption) variants and their accompanying classification scheme, the per-scale metadata, and the spatial footprint layer are the original work of the authors and are licensed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license. Use for any purpose — including the training and evaluation of machine-learning models, and commercial use — is permitted provided appropriate attribution is given. For machine-learning use, attribution is satisfied by citing this dataset and its accompanying paper at the corpus level; per-sample attribution is not required. Underlying Copernicus data. The Sentinel-2 imagery and the CORINE Land Cover 2018 (CLC2018) reference are provided under the Copernicus open data policy (Regulation (EU) 2021/696 establishing the Union Space Programme, together with Commission Delegated Regulation (EU) No 1159/2013 on the Copernicus data and information policy), which permits free use, reproduction, adaptation, and redistribution, including for commercial purposes. The underlying Sentinel-2 and CLC2018 data remain the sole property of the European Union. The following attributions apply and must be retained in any redistribution or derivative work: Sentinel-2 imagery: "Contains modified Copernicus Sentinel data 2018." CORINE Land Cover 2018: "© European Union, Copernicus Land Monitoring Service 2018, European Environment Agency (EEA)." The Copernicus data were modified by the authors (resampling to a 10 m grid, multi-scale patch extraction, and derivation of land-cover labels). Neither the European Union, the European Environment Agency, nor the Copernicus programme endorses this dataset or any conclusions drawn from it. Code: https://github.com/geosense-lab/SENSERO



