TestDS
收藏资源简介:
GeoMeld 是一个大规模多模态遥感数据集,专为语义基础建模设计。该数据集包含约250万个空间对齐的样本,覆盖多种传感模态和空间分辨率,并配有通过代理流程生成的语义基础描述。每个样本构成一个跨多分辨率空间对齐的多模态元组: 1. 高分辨率(约1米):覆盖美国本土的1米地面采样距离(GSD)RGB影像,来自国家农业影像计划(NAIP),并与低分辨率卫星模态共配准。 2. 中分辨率(10米,标准化网格):包括Sentinel-2多光谱光学影像(12波段)、Sentinel-1 SAR后向散射(VV、VH、HH、HV)、ASTER-DEM高程和地形坡度、冠层高度以及土地覆盖产品(Dynamic World、ESA WorldCover)。 所有10米模态对齐到128×128网格,而高分辨率NAIP影像提供1280×1280的细粒度空间上下文。数据集以约50GB的WebDataset(.tar)分片形式归档,支持直接从Hugging Face流式传输到PyTorch训练管道。数据集包含两个子集:NAIP(高分辨率)和非NAIP(中分辨率),可通过文件名后缀区分。每个样本还包括地理元数据(位置、区域描述符)和语义基础描述。 该数据集适用于图像分类、图像分割、零样本图像分类、文本到图像、图像到文本和特征提取等任务。
GeoMeld is a large-scale multimodal remote sensing dataset designed explicitly for semantic grounding modeling. It contains approximately 2.5 million spatially aligned samples, covering multiple sensing modalities and spatial resolutions, and is accompanied by semantic grounding descriptions generated via proxy workflows. Each sample forms a spatially aligned multimodal tuple across multiple resolutions: 1. High-resolution (~1 meter): 1-meter ground sampling distance (GSD) RGB imagery covering the contiguous United States, sourced from the National Agriculture Imagery Program (NAIP), and co-registered with low-resolution satellite modalities. 2. Medium-resolution (10 meters, standardized grid): includes Sentinel-2 multispectral optical imagery (12 bands), Sentinel-1 SAR backscatter (VV, VH, HH, HV), ASTER-DEM elevation and terrain slope, canopy height, and land cover products (Dynamic World, ESA WorldCover). All 10-meter modalities are aligned to a 128×128 grid, while the high-resolution NAIP imagery provides fine-grained spatial context at 1280×1280. The dataset is archived as ~50 GB WebDataset (.tar) shards, supporting direct streaming from Hugging Face into PyTorch training pipelines. It includes two subsets: NAIP (high-resolution) and non-NAIP (medium-resolution), which can be distinguished via filename suffixes. Each sample also incorporates geographic metadata (location, region descriptors) and semantic grounding descriptions. This dataset is applicable to tasks including image classification, image segmentation, zero-shot image classification, text-to-image generation, image-to-text generation, and feature extraction.




