遇见数据集

Data from: Minimum required number of specimen records to develop accurate species distribution models

收藏
DataONE2015-06-01 更新2024-06-27 收录
数据链接:
官方服务:

资源简介:

Species Distribution Models (SDMs) are widely used to predict the occurrence of species. Because SDMs generally use presence-only data, validation of the predicted distribution and assessing model accuracy is challenging. Model performance depends on both sample size and species’ prevalence, being the fraction of the study area occupied by the species. Here, we present a novel method using simulated species to identify the minimum number of records required to generate accurate SDMs for taxa of different pre-defined prevalence classes. We quantified model performance as a function of sample size and prevalence and found model performance to increase with increasing sample size under constant prevalence, and to decrease with increasing prevalence under constant sample size. The Area Under the Curve (AUC) is commonly used as a measure of model performance. However, when applied to presence-only data it is prevalence-dependent and hence not an accurate performance index. Testing the AUC of an SDM for significant deviation from random performance provides a good alternative. We assessed the minimum number of records required to obtain good model performance for species of different prevalence classes in a virtual study area and in a real African study area. The lower limit depends on the species’ prevalence with absolute minimum sample sizes as low as 3 for narrow-ranged and 13 for widespread species for our virtual study area which represents an ideal, balanced, orthogonal world. The lower limit of 3, however, is flawed by statistical artefacts related to modelling species with a prevalence below 0.1. In our African study area lower limits are higher, ranging from 14 for narrow-ranged to 25 for widespread species. We advocate identifying the minimum sample size for any species distribution modelling by applying the novel method presented here, which is applicable to any taxonomic clade or group, study area or climate scenario.

物种分布模型(Species Distribution Models, SDMs)被广泛用于预测物种的发生情况。由于SDMs通常仅使用仅存在数据,因此对预测分布进行验证以及评估模型精度颇具挑战。模型性能同时取决于样本量与物种流行率——即研究区域内被该物种占据的面积比例。本文提出一种全新方法,借助模拟物种来确定针对不同预设流行率类群构建精准SDMs所需的最小记录数量。我们将模型性能量化为样本量与流行率的函数,结果发现:在流行率固定的情况下,模型性能随样本量增加而提升;在样本量固定的情况下,模型性能随流行率升高而下降。曲线下面积(Area Under the Curve, AUC)常被用作模型性能的评价指标。但当应用于仅存在数据时,AUC受流行率影响,因此并非准确的性能评价指标。对SDM的AUC进行显著性检验,判断其与随机性能是否存在显著差异,是一种良好的替代方案。我们在虚拟研究区域与真实非洲研究区域中,评估了不同流行率类群的物种获得良好模型性能所需的最小记录数量。在代表理想、均衡且正交环境的虚拟研究区域中,最小样本量下限取决于物种流行率:狭域分布物种的绝对最小样本量低至3,广布物种则为13。但当流行率低于0.1时,这一下限(3)会受到与建模相关的统计假象影响而存在缺陷。在我们的非洲研究区域中,最小样本量下限更高,范围为狭域分布物种14至广布物种25。我们倡议,针对任何物种分布建模工作,均应通过本文提出的全新方法确定最小样本量;该方法适用于任意分类类群、研究区域或气候情景。

创建时间:
2015-06-01
二维码
社区交流群
二维码
科研交流群
商业服务