PNV - Probability distribution for Pinus halepensis
收藏资源简介:
Overview: Potential Natural Vegetation (PNV): potential probability of occurrence for the Aleppo pine from 2018 to 2020 Traceability (lineage): This is an original dataset produced with a machine learning framework which used a combination of point datasets and raster datasets as inputs. Point dataset is a harmonized collection of tree occurrence data, comprising observations from National Forest Inventories (EU-Forest), GBIF and LUCAS. The complete dataset is available on Zenodo. Raster datasets used as input are: monthly time series air and surface temperature and precipitation from a reprocessed version of the Copernicus ERA5 dataset; long term averages of bioclimatic variables from CHELSA; elevation, slope and other elevation-derived metrics and long term monthly averages snow probability. For a more comprehensive list refer to Bonannella et al. (2022) (in review, preprint available at: https://doi.org/10.21203/rs.3.rs-1252972/v1). Scientific methodology: Probability and uncertainty maps were the output of a spatiotemporal ensemble machine learning framework based on stacked regularization. Three base models (random forest, gradient boosted trees and generalized linear models) were first trained on the input dataset and their predictions were used to train an additional model (logistic regression) which provided the final predictions. More details on the whole workflow are available in the listed publication. Usability: Probability maps are particularly useful when compared with existing products of potential distribution of species or when combined with maps of realized distribution: gaps in potential and realized distribution can be identified and used as information for future programs of tree planting or forest restoration. Uncertainty quantification: Uncertainty is quantified by taking the standard deviation of the probabilities predicted by the three components of the spatiotemporal ensemble model. Data validation approaches: Distribution maps were validated using a spatial 5-fold cross validation following the workflow detailed in the listed publication. Completeness: The raster files perfectly cover the entire Geo-harmonizer region as defined by the landmask raster dataset available here. Consistency: Areas which are outside of the calibration area of the point dataset (Iceland, Norway) usually have high uncertainty values. This is not only a problem of extrapolation but also of poor representation in the feature space available to the model of the conditions that are present in this countries. Positional accuracy: The rasters have a spatial resolution of 30m. Temporal accuracy: The maps cover the period 2018 - 2020 Thematic accuracy: Both probability and uncertainty maps contain values from 0 to 100: in the case of probability maps, they indicate the probability of occurrence of a single individual of the target species, while uncertainty maps indicate the standard deviation of the ensemble model.
Overview: 潜在自然植被(Potential Natural Vegetation, PNV):2018年至2020年阿勒颇松的潜在发生概率。 Traceability (lineage): 本数据集为原创数据集,采用结合点数据集与栅格数据集作为输入的机器学习框架生成。点数据集为经统一协调的树木出现数据合集,包含来自国家森林清查(EU-Forest)、全球生物多样性信息设施(GBIF)以及LUCAS的观测数据。完整数据集可在Zenodo平台获取。作为输入的栅格数据集包括:经再处理的哥白尼ERA5(Copernicus ERA5)数据集的逐月气温、地表温度与降水时间序列;CHELSA数据集的生物气候变量长期平均值;高程、坡度及其他高程衍生指标,以及逐月降雪概率长期平均值。完整数据集清单可参考Bonannella等人(2022)的研究(已投稿待审,预印本链接:https://doi.org/10.21203/rs.3.rs-1252972/v1)。 Scientific methodology: 概率与不确定性地图为基于堆叠正则化的时空集成机器学习框架的输出成果。首先基于输入数据集训练三个基础模型(随机森林、梯度提升树与广义线性模型),随后以这三个模型的预测结果作为输入,训练额外的逻辑回归模型以生成最终预测结果。完整工作流的更多细节可参见上述已列出的研究文献。 Usability: 概率地图与现有物种潜在分布产品对比,或与实际分布地图结合使用时尤为实用:可识别潜在分布与实际分布之间的空白区域,为未来植树造林或森林恢复项目提供决策依据。 Uncertainty quantification: 不确定性通过计算时空集成模型三个组成部分的预测概率的标准差来量化。 Data validation approaches: 遵循上述文献中详述的工作流,采用空间5折交叉验证对分布地图进行验证。 Completeness: 栅格文件完全覆盖由此处提供的陆面mask栅格数据集所定义的地理协调区域。 Consistency: 位于点数据集校准区域之外的区域(冰岛、挪威)通常具有较高的不确定性值。这不仅是外推导致的问题,同时也是模型特征空间未能充分覆盖这些国家的立地条件所致。 Positional accuracy: 栅格数据的空间分辨率为30米。 Temporal accuracy: 该地图覆盖2018年至2020年的时间段。 Thematic accuracy: 概率地图与不确定性地图的取值范围均为0至100:其中概率地图的数值代表目标物种单个个体的出现概率,不确定性地图的数值则代表集成模型的标准差。



