遇见数据集

Core-S2RGB-249k-TIPSv2

收藏
魔搭社区2026-08-23 更新2026-08-16 收录
官方服务:

资源简介:

# Core-S2RGB-249k-TIPSv2 Geospatial-vision embedding dataset computed from **Core-S2RGB-249k** using the **TIPSv2** model. ## Overview | Property | Value | |----------|-------| | Source imagery | Core-S2RGB-249k (248,719 patches) | | Model | TIPSv2-B/14 ([google/tipsv2-b14](https://huggingface.co/google/tipsv2-b14)) | | Input bands | RGB [B04, B03, B02] | | Embedding dimension | 768 | | Output format | GeoParquet | | License | CC-BY-SA-4.0 | ## Computation Pipeline 1. **Pre-processing**: Each Sentinel-2 L2A patch is read from the source parquet files. The RGB bands [B04, B03, B02] are stacked and scaled by `2.5 × DN / 1e4`, clipped to [0, 1], to convert digital numbers to display-range reflectance. 2. **Resize**: The RGB tensor is interpolated to the TIPSv2-B/14 input size of **448 × 448** pixels. 3. **Encoding**: The tensor is fed into the TIPSv2 image encoder (a ViT-B/14 trained with spatially-aware vision-language pretraining, enhancing patch-text alignment) to extract a 768-dimensional image embedding. 4. **Post-processing**: Embeddings are L2-normalised by the model encoder, so a dot product between rows equals cosine similarity. 5. **Geospatial metadata**: The original UTM footprint is reprojected to EPSG:4326 (WGS-84) to obtain the `geometry`, `centre_lat`, and `centre_lon` fields. Additional metadata (`product_id`, `grid_cell`, `timestamp`, `utm_crs`, `pixel_bbox`) is preserved from the source dataset. ## File Layout ``` . ├── TIPSv2_b14_crop_448x448.parquet # Main embedding GeoParquet (248,719 rows) └── README.md ``` ## Schema | Column | Type | Description | |--------|------|-------------| | `unique_id` | string | SHA-256 hash of geometry + timestamp + product_id + embedding | | `embedding` | list<float> | 768-dim TIPSv2 feature vector | | `timestamp` | datetime | Acquisition timestamp | | `product_id` | string | Original Sentinel-2 product identifier | | `grid_cell` | string | Major-TOM grid cell identifier | | `grid_row_u` | int16 | Grid row index | | `grid_col_r` | int16 | Grid column index | | `geometry` | geometry | WGS-84 polygon (footprint) | | `centre_lat` | float32 | Latitude of patch centre | | `centre_lon` | float32 | Longitude of patch centre | | `utm_footprint` | string | Original UTM footprint as WKT | | `utm_crs` | string | UTM CRS (e.g. EPSG:32633) | | `pixel_bbox` | list<int> | Pixel bounding box [x_min, y_min, x_max, y_max] | | `parquet_url` | string | Source parquet file path in the image dataset | | `parquet_row` | int64 | Row index within the source parquet file | ## Usage ```python import pandas as pd df = pd.read_parquet("TIPSv2_b14_crop_448x448.parquet") print(len(df), "embeddings") print(df.iloc[0].embedding.shape) # (768,) ``` ## Citation If you use this embedding dataset, please cite the original Major-TOM paper and the TIPSv2 paper: ```bibtex @article{cao2026tipsv2, title={TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment}, author={Cao, Bingyi and Chen, Koert and Maninis, Kevis-Kokitsi and Chen, Kaifeng and Karpur, Arjun and Xia, Ye and Dua, Sahil and Dabral, Tanmaya and Han, Guangxing and Han, Bohyung and Ainslie, Joshua and Bewley, Alex and Jacob, Mithun and Wagner, Ren{\'e} and Ramos, Washington and Choromanski, Krzysztof and Seyedhosseini, Mojtaba and Zhou, Howard and Araujo, Andr{\'e}}, journal={arXiv preprint arXiv:2604.12012}, year={2026} } ``` ```bibtex @inproceedings{francis2024majortom, title={Major TOM: Expandable Datasets for Earth Observation}, author={Francis, Alistair and Czerkawski, Mikolaj}, year={2024}, booktitle={IGARSS 2024}, eprint={2402.12095}, archivePrefix={arXiv} } ```

提供机构:
maas
创建时间:
2026-08-09
二维码
社区交流群
二维码
科研交流群
商业服务