遇见数据集

The importance of the sampling design in mapping woody cover in arid ecosystems

收藏
Zenodo2026-01-13 更新2026-05-26 收录
官方服务:

资源简介:

Accurate calibration data remains a major constraint in ecological remote sensing, particularly in arid ecosystems where sparse and heterogeneous woody vegetation is challenging to detect. Despite advances in sensor technology and modelling algorithms, optimal calibration sampling design has received limited attention, with current approaches varying unsystematically (35–1,000 plots/1,000 km²) without empirical benchmarks. Here, we develop and evaluate a quantitative framework for optimizing sampling strategies in remote-sensing models of woody cover through systematic comparison of (i) calibration data source (field surveys versus photointerpretation), (ii) spatial configuration (clustered versus dispersed), and (iii) sampling density. This integrated approach enables isolating the relative contributions of each design component to mapping accuracy—a critical gap in current remote sensing methodology. We applied this framework to Sentinel-1, Sentinel-2, combined Sentinel-1+2, and AlphaEarth Foundations across Madagascar's arid southwest, validating predictions through spatial cross-validation using independent field plots. Photointerpretation substantially outperformed field-based calibration under typical arid-zone constraints (R²=0.88, RMSE=0.11 versus R²=0.46–0.66, RMSE=0.17–0.21). Performance saturated at 20.7 - 41.4 dispersed calibration plots per 1,000 km² across all predictors, beyond which gains became marginal. Dispersed strategies required half as many samples as clustered designs to achieve comparable accuracy, demonstrating that spatial distribution outweighs sample size. Once adequate sampling density and distribution were achieved, Sentinel-1+2 and AlphaEarth performed similarly (R²≈0.85–0.86), indicating that sampling design outweighs predictor complexity. Our framework provides empirically derived operational thresholds (minimum: 15.5 plots/1,000 km²; optimal: 20.7 -41.4 plots/1,000 km²) and a transferable methodology for determining optimal calibration densities in heterogeneous ecosystems where logistical constraints limit field sampling.

精准定标数据仍是生态遥感领域的一大制约瓶颈,在干旱生态系统中尤为凸显——这类区域内稀疏且异质的木本植被难以被精准探测。尽管传感器技术与建模算法已取得显著进展,但最优定标采样设计的相关研究仍较为匮乏,当前的采样方案多为非系统性的多样化尝试(采样密度范围为35~1000个样地/1000平方千米),且缺乏经验基准。为此,本研究构建并评估了一套定量框架,用于优化木本植被盖度遥感反演模型的采样策略,通过系统性对比三类要素:(i) 定标数据来源(野外调查与目视解译)、(ii) 空间配置(聚类式与分散式)以及(iii) 采样密度。该整合方法可实现各设计组分对制图精度相对贡献的分离,这正是当前遥感方法学中亟待填补的关键空白。 我们将该框架应用于马达加斯加西南部干旱区域的哨兵1号(Sentinel-1)、哨兵2号(Sentinel-2)、哨兵1号+2号融合数据以及AlphaEarth Foundations遥感数据集,通过独立野外样地开展空间交叉验证以验证模型预测结果。在典型干旱区域约束条件下,目视解译定标数据的表现显著优于野外调查定标(决定系数R²=0.88,均方根误差RMSE=0.11;而野外调查定标的R²为0.46~0.66,RMSE为0.17~0.21)。在所有预测变量下,当采样密度达到每1000平方千米20.7~41.4个分散样地时,模型性能趋于饱和,超过该密度后精度提升幅度微乎其微。分散式采样方案仅需聚类式采样一半的样本量即可达到相当的精度,这表明空间分布格局相较于采样规模对精度的影响更为显著。当采样密度与空间分布达到最优水平后,哨兵1号+2号融合数据与AlphaEarth Foundations的表现相近(R²≈0.85~0.86),这说明采样设计相较于预测变量复杂度对精度的影响更为关键。 本研究框架提供了经验推导得到的可操作阈值(最低阈值:15.5个样地/1000平方千米;最优阈值:20.7~41.4个样地/1000平方千米),并为后勤约束限制野外采样的异质生态系统提供了可迁移的最优定标密度确定方法。

提供机构:
Zenodo
创建时间:
2026-01-13
二维码
社区交流群
二维码
科研交流群
商业服务