Core-S2L2A-249k-SatCLIP
收藏资源简介:
# Core-S2RGB-249k-SatCLIP Geospatial-vision embedding dataset computed from **Core-S2L2A-249k** using the **SatCLIP** model. ## Overview | Property | Value | |----------|-------| | Source imagery | Core-S2L2A-249k (248,719 patches, 384 × 384 px) | | Model | SatCLIP (ResNet-50 + location encoder) | | Input bands | All 12 Sentinel-2 bands [B01, B02, B03, B04, B05, B06, B07, B08, B8A, B09, B11, B12] | | Embedding dimension | 512 | | Output format | GeoParquet | | License | CC-BY-SA-4.0 | ## Computation Pipeline 1. **Pre-processing**: Each 384 × 384 Sentinel-2 L2A patch is read from the source parquet files. All 12 spectral bands are stacked and normalised by dividing by `1e4` to convert digital numbers to reflectance. 2. **Resize & Pad**: The 12-band tensor is interpolated to the SatCLIP input size of **224 × 224** pixels. Because SatCLIP expects 13 input channels, a zero-filled B10 channel is padded at index 10. 3. **Encoding**: The 13-channel tensor is fed into the SatCLIP image encoder (a ResNet-50 trained with location-aware contrastive learning on satellite imagery) to extract a 512-dimensional image embedding. 4. **Post-processing**: No L2-normalisation is applied during dataset generation; normalisation is performed at retrieval time if required. 5. **Geospatial metadata**: The original UTM footprint is reprojected to EPSG:4326 (WGS-84) to obtain the `geometry`, `centre_lat`, and `centre_lon` fields. Additional metadata (`product_id`, `grid_cell`, `timestamp`, `utm_crs`, `pixel_bbox`) is preserved from the source dataset. ## File Layout ``` . ├── SatCLIP_crop_384x384.parquet # Main embedding GeoParquet (248,719 rows) └── README.md ``` ## Schema | Column | Type | Description | |--------|------|-------------| | `unique_id` | string | SHA-256 hash of geometry + timestamp + product_id + embedding | | `embedding` | list<float> | 512-dim SatCLIP feature vector | | `timestamp` | datetime | Acquisition timestamp | | `product_id` | string | Original Sentinel-2 product identifier | | `grid_cell` | string | Major-TOM grid cell identifier | | `grid_row_u` | int16 | Grid row index | | `grid_col_r` | int16 | Grid column index | | `geometry` | geometry | WGS-84 polygon (footprint) | | `centre_lat` | float32 | Latitude of patch centre | | `centre_lon` | float32 | Longitude of patch centre | | `utm_footprint` | string | Original UTM footprint as WKT | | `utm_crs` | string | UTM CRS (e.g. EPSG:32633) | | `pixel_bbox` | list<int> | Pixel bounding box [x_min, y_min, x_max, y_max] | | `parquet_url` | string | Source parquet file path in the image dataset | | `parquet_row` | int64 | Row index within the source parquet file | ## Usage ```python import pandas as pd df = pd.read_parquet("SatCLIP_crop_384x384.parquet") print(len(df), "embeddings") print(df.iloc[0].embedding.shape) # (512,) ``` ## Citation If you use this embedding dataset, please cite the original Major-TOM paper and the SatCLIP paper: ```bibtex @article{zheng2026earthembeddingexplorer, title={EarthEmbeddingExplorer: A Web Application for Cross-Modal Retrieval of Global Satellite Images}, author={Zheng, Yijie and Wu, Weijie and Wu, Bingyue and Zhao, Long and Li, Guoqing and Czerkawski, Mikolaj and Klemmer, Konstantin}, journal={arXiv preprint arXiv:2603.29441}, year={2026}, note={ICLR 2026 Workshop ML4RS Tutorial Track (oral)} } ``` ```bibtex @inproceedings{francis2024majortom, title={Major TOM: Expandable Datasets for Earth Observation}, author={Francis, Alistair and Czerkawski, Mikolaj}, year={2024}, booktitle={IGARSS 2024}, eprint={2402.12095}, archivePrefix={arXiv} } ```



