Major-TOM/Core-S2RGB-249k-SigLIP
收藏资源简介:
Core-S2RGB-249k-SigLIP是一个基于Core-S2L2A-249k数据集并使用SigLIP模型计算的视觉语言嵌入数据集。数据集包含248,719个384×384像素的Sentinel-2 L2A补丁,通过RGB波段(B04, B03, B02)生成真彩色图像,并使用SigLIP视觉编码器提取1152维图像嵌入。数据集还保留了地理空间元数据,如几何形状、中心经纬度等。数据集以GeoParquet格式存储,包含嵌入向量、时间戳、产品标识符、网格单元标识符、地理坐标等信息。
Core-S2RGB-249k-SigLIP is a vision-language embedding dataset based on the Core-S2L2A-249k dataset, with its vision-language embeddings computed using the SigLIP model. The dataset contains 248,719 Sentinel-2 L2A patches with a resolution of 384×384 pixels. True-color images are generated from these patches via the RGB bands (B04, B03, B02), and 1152-dimensional image embeddings are extracted using the SigLIP vision encoder. The dataset also retains geospatial metadata such as geometry, central longitude and latitude, and other related information. It is stored in GeoParquet format, containing embedding vectors, timestamps, product identifiers, grid cell identifiers, geographic coordinates, and other relevant information.




