遇见数据集

Contiguous U.S. Daily PM<sub>2.5</sub> Measurements (2021) – Benchmark Dataset for Location‑Encoder Evaluation

收藏
DataCite Commons2025-05-23 更新2025-09-08 收录
官方服务:

资源简介:

Deep learning models have demonstrated success in geospatial applications, yet quantifying the role of geolocation information in enhancing model performance and geographic generalizability remains underexplored. A new generation of location encoders have emerged with the goal of capturing attributes present at any given location for downstream use in predictive modeling. Being a nascent area of research, their evaluation has remained largely limited to static tasks such as species distributions or average temperature mapping. In this paper, we discuss and quantify the impact of incorporating geolocation into deep learning for a real-world application domain that is characteristically dynamic (with fast temporal change) and spatially heterogeneous at high resolutions: estimating surface-level daily PM<sub>2.5</sub> levels using remotely sensed and ground-level data. We build on a recently published deep learning-based PM<sub>2.5</sub> estimation model that achieves state-of-the-art performance on data observed in the contiguous United States. We examine three approaches for incorporating geolocation: excluding geolocation as a baseline, using raw geographic coordinates, and leveraging pretrained location encoders. We evaluate each approach under within-region (WR) and out-of-region (OoR) evaluation scenarios. Aggregate performance metrics indicate that while naïve incorporation of raw geographic coordinates improves within-region performance by retaining the interpolative value of geographic location, it can hinder generalizability across regions. In contrast, pretrained location encoders like GeoCLIP enhance predictive performance and geographic generalizability for both WR and OoR scenarios. However, our qualitative analysis reveals artifact patterns caused by high-degree basis functions and sparse upstream samples in certain areas, and our ablation results indicate varying performance among location encoders such as SatCLIP vs. GeoCLIP. To the best of our knowledge, this is a first integration and systematic evaluation of location encoders in a complex, temporally dynamic estimation scenario. In addition to guiding better model development for air pollution estimation and location encoders, this study provides insights for effective incorporation of location into deep learning for geospatial predictive tasks.

深度学习模型已在地理空间应用领域展现出卓越成效,但量化地理位置信息对提升模型性能与地理泛化性的作用,仍未得到充分探索。新一代位置编码器(location encoders)应运而生,其目标是捕捉任意给定位置的相关属性,以供下游预测建模任务使用。作为新兴研究方向,此类编码器的评估目前大多局限于物种分布、平均气温制图等静态任务。本文针对一个兼具快速时间变化与高分辨率空间异质性特征的动态现实应用场景——利用遥感数据与地面观测数据估算地表每日PM₂.₅浓度——探讨并量化了将地理位置信息融入深度学习模型的影响。我们基于近期公开的一款基于深度学习的PM₂.₅估算模型展开研究,该模型在美国本土区域的观测数据上实现了当前最优(state-of-the-art)性能。我们研究了三种融入地理位置信息的方案:将地理位置作为基线特征予以排除、使用原始地理坐标,以及利用预训练位置编码器。我们在区域内(within-region, WR)与跨区域(out-of-region, OoR)两种评估场景下,对每种方案开展评估。综合性能指标显示,尽管直接使用原始地理坐标的融入方式,通过保留地理位置的插值价值提升了区域内模型性能,但却会阻碍模型的跨区域泛化能力。与之形成鲜明对比的是,GeoCLIP这类预训练位置编码器,可同时提升区域内与跨区域场景下的预测性能与地理泛化性。不过,我们的定性分析揭示了由高次基函数与部分区域稀疏的上游采样数据所导致的伪影模式;同时消融实验结果表明,不同位置编码器(如SatCLIP与GeoCLIP)的性能存在差异。据我们所知,本文是首次在复杂的动态时间估算场景中,对位置编码器进行集成与系统性评估的研究。本研究不仅为空气污染估算模型与位置编码器的优化开发提供了指导,也为在地理空间预测任务中高效融入位置信息的深度学习方案提供了理论见解。

提供机构:
figshare
创建时间:
2025-05-23
搜集汇总
数据集介绍
Contiguous U.S. Daily PM<sub>2.5</sub> Measurements (2021) – Benchmark Dataset for Location‑Encoder Evaluation 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务