遇见数据集

Data and materials for "Climate limits the niche, settlement mosaics expand it: implications for the Asian range expansion of invasive Spanish slug Arion vulgaris"

收藏
Zenodo2026-02-28 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains the data and supplementary materials associated with a manuscript currently in preparation, provisionally entitled “Climate limits the niche, settlement mosaics expand it: implications for the Asian range expansion of invasive Spanish slug Arion vulgaris”. The final title and bibliographic details will be updated upon acceptance. The materials include two major components:(1) species distribution modelling outputs, including environmental significance tables, AUC summaries, permutation importance heatmaps, partial dependence analyses, and global projections; and(2) phylogenetic data supporting species identification, including the COI alignment and resulting maximum-likelihood and Bayesian trees. All modelling scripts used to generate these results are openly available on GitHub (see Related identifiers). Data and results deposited here allow full reproduction of the analyses presented in the manuscript. Repository structure Description of Supplementary Tables This archive contains unified performance metrics and paired comparisons among predictor configurations for all modelling algorithms. 1. all_buffers_METRICS_standard_all_runs.csv Complete table of model evaluation metrics extracted for all predictor configurations, buffer sizes, and algorithms. Each row corresponds to a single model run or ensemble model and includes: scenario – predictor set (Climate or Urb+Clim) buffer_km – pseudo absence buffer distance in kilometres buffer_label – formatted configuration label is_block_shuffle – whether anthropogenic layers were spatially shuffled using 300 km blocks source – single_models or ensemble_models algo – modelling algorithm (GBM, GLM, MAXENT, RF) metric_std – standardized metric (AUC, TSS, LOOIC) metric_value – metric value (LOOIC normalized per validation point) additional diagnostic columns including calibration, validation, and evaluation values This file represents the complete non aggregated dataset of model performance. 2. all_buffers_METRICS_standard_summary_by_algo.csv Summary statistics of model performance aggregated by buffer configuration, algorithm, and metric. For each combination, the table provides: number of model runs mean and standard deviation of metric values number of calibration and validation records This table is used for descriptive comparison among predictor configurations. 3. all_buffers_METRICS_standard_summary_by_buffer.csv Performance metrics aggregated at the buffer configuration level across algorithms. This file provides overall mean values and variability for each predictor scenario and buffer size. 4. all_buffers_LOOIC_proxy_all_runs.csv Per run LOOIC proxy values calculated from out of sample predictions. LOOIC values are computed from validation log likelihoods and normalized per validation observation to allow comparison across runs with different validation sample sizes. 5. all_buffers_delta_vs_25km_by_run.csv Paired differences in performance relative to the baseline model Clim+Antr 25 km. Each row represents a paired comparison for a given pseudo absence set, run, algorithm, and metric. The column delta_metric represents the difference between the given configuration and the baseline model. For AUC and TSS, positive values indicate improvement relative to baseline.For LOOIC normalized per point, negative values indicate improvement. This table is used for paired statistical comparisons among configurations. 6. all_buffers_delta_vs_25km_summary_by_algo.csv Mean and standard deviation of paired differences aggregated by predictor configuration, algorithm, and metric. This table includes: number of paired comparisons mean delta standard deviation of delta These values correspond to those visualized in the performance difference figures. Reproducibility All metrics were extracted directly from BIOMOD2 model objects.Block shuffle configurations represent the mean performance across 100 spatially shuffled replicates.LOOIC values are calculated as validation based log likelihood proxies and normalized per validation observation. 7. environmental_significance.xlsx Relative contributions of predictors in ensemble models 8. Nekhaev et al Suppl Figs.pdf File containing supplementary figures including original computer-generated versions and models projections 9. Folder containing model projections generated using climate-only predictors as well as combined climate and land cover predictors, based on models fitted with different pseudo-absence buffer distances. The folder also includes a QGIS project file with all layers pre-configured for visualization. Methods summary Molecular techniques and phylogenetic analysis.DNA was extracted from foot tissues of two specimens, and COI was amplified using standard Folmer primers. Bidirectional Sanger sequencing was performed in Almaty. Reads were assembled in Ugene; final sequences (484 bp) are deposited in GenBank. The alignment was produced with MAFFT and filtered with gBlocks. The optimal substitution model (TPM3u+G+I) was selected using phangorn in R. Maximum-likelihood analysis with 100,000 bootstrap replicates was conducted in R, and Bayesian inference was run in RevBayes with three independent MCMC runs. Convergence diagnostics were performed in R. All phylogeny scripts are available on GitHub. Occurrence data.Distribution modelling used GBIF records for Arion vulgaris, A. rufus and A. ater, supplemented by visually verified iNaturalist observations from Asia. Due to unreliable external morphological separation and strong ecological niche overlap among these taxa, all records were treated as a single species complex. Occurrences were merged, visually checked, and aggregated into 25 × 25 km grid cells to reduce spatial autocorrelation, resulting in 2757 final presence points. Environmental predictors.We used two predictor groups: (1) bioclimatic variables from WorldClim 2.1, reduced to five weakly correlated predictors (Bio1, Bio2, Bio7, Bio12, Bio14), and (2) land-cover fractions from the CDR and Sentinel-3 CCI Land Cover datasets (2022), aggregated into intensive agriculture, mosaic agro-natural landscapes and urban areas. Environmental layers were processed in R using terra and geodata. Species distribution modelling.Global niche models for the Arion vulgaris–A. rufus–A. ater complex were built using biomod2 in R at 2.5 arc-minute resolution and projected onto Asia at 0.5 arc minutes. Four algorithms were used (GLM, GBM, RF, Maxent) with 12,000 pseudo-absences sampled outside a 200 km buffer around occurrences. Each model was trained with ten 70/30 cross-validation splits, and evaluated using AUC. An ensemble prediction was constructed using accuracy-weighted means. All modelling scripts are available on GitHub. Regional predictor assessment.To analyse localised effects of climatic and anthropogenic factors, the Asian study region was subdivided into 12 subregions (3 latitudinal × 4 longitudinal bands) with a European control area representing the native range. For each subregion, permutation importance (PI) of predictors and partial dependence (PD) curves were computed. PD curves were based on 100,000 random pixels per region. Analyses were performed in R (biomod2), and visualisations were produced with ggplot2.

提供机构:
Zenodo
创建时间:
2026-02-28
二维码
社区交流群
二维码
科研交流群
商业服务