Using Historical Atlas Data to Develop High-Resolution Distribution Models of Freshwater Fishes
收藏资源简介:
Understanding the spatial pattern of species distributions is fundamental in biogeography, and conservation and resource management applications. Most species distribution models (SDMs) require or prefer species presence and absence data for adequate estimation of model parameters. However, observations with unreliable or unreported species absences dominate and limit the implementation of SDMs. Presence-only models generally yield less accurate predictions of species distribution, and make it difficult to incorporate spatial autocorrelation. The availability of large amounts of historical presence records for freshwater fishes of the United States provides an opportunity for deriving reliable absences from data reported as presence-only, when sampling was predominantly community-based. In this study, we used boosted regression trees (BRT), logistic regression, and MaxEnt models to assess the performance of a historical metacommunity database with inferred absences, for modeling fish distributions, investigating the effect of model choice and data properties thereby. With models of the distribution of 76 native, non-game fish species of varied traits and rarity attributes in four river basins across the United States, we show that model accuracy depends on data quality (e.g., sample size, location precision), species’ rarity, statistical modeling technique, and consideration of spatial autocorrelation. The cross-validation area under the receiver-operating-characteristic curve (AUC) tended to be high in the spatial presence-absence models at the highest level of resolution for species with large geographic ranges and small local populations. Prevalence affected training but not validation AUC. The key habitat predictors identified and the fish-habitat relationships evaluated through partial dependence plots corroborated most previous studies. The community-based SDM framework broadens our capability to model species distributions by innovatively removing the constraint of lack of species absence data, thus providing a robust prediction of distribution for stream fishes in other regions where historical data exist, and for other taxa (e.g., benthic macroinvertebrates, birds) usually observed by community-based sampling designs.
明晰物种分布的空间格局,是生物地理学、物种保护与资源管理应用的核心基础。多数物种分布模型(Species Distribution Models, SDMs)为精准估算模型参数,通常需要或优先获取物种的出现与未出现数据。然而,当前主流观测数据多为不可靠或未报告的物种未出现记录,这极大限制了物种分布模型的应用。仅基于出现记录的模型通常生成的物种分布预测精度较低,且难以纳入空间自相关(spatial autocorrelation)效应。美国淡水鱼类存在大量历史出现记录,且其采样多以群落调查为核心,这为从仅标注出现的记录中推导可靠的未出现数据提供了契机。本研究采用提升回归树(Boosted Regression Trees, BRT)、逻辑回归与最大熵模型(MaxEnt),针对基于推断未出现数据构建的历史群落元数据库,开展鱼类分布建模性能评估,以此探究模型选择与数据属性对建模结果的影响。本研究以美国4个流域内76种具有不同功能性状与稀有性特征的本土非游钓鱼类为研究对象,构建其分布模型,结果表明模型精度受数据质量(如样本量、点位定位精度)、物种稀有性、统计建模方法以及空间自相关考量因素的综合影响。受试者工作特征曲线下面积(Area Under the Receiver Operating Characteristic Curve, AUC)的交叉验证结果显示,针对地理分布范围广但本地种群规模小的物种,高分辨率空间出现-未出现模型的AUC值普遍较高。物种出现率会影响模型训练阶段的表现,但对验证阶段的AUC值无显著影响。通过偏依赖图识别的关键生境预测因子,以及解析得到的鱼类与生境的关联关系,与多数已有研究结果一致。基于群落调查的物种分布模型框架,通过创新性地突破物种未出现数据缺失的限制,拓展了物种分布建模的能力边界;该框架可为存在历史调查数据的其他区域的溪流鱼类,以及通常采用群落采样方案的其他生物类群(如底栖大型无脊椎动物、鸟类)提供可靠的分布预测结果。



