遇见数据集

Data from: Disentangling spatial and environmental sampling bias in occurrence data: A species distribution model for the mosquito Aedes vexans

收藏
Zenodo2026-04-14 更新2026-05-26 收录
官方服务:

资源简介:

File Description Model Types In this study, we developed and compared several species distribution models (SDMs) for Aedes vexans, each differing in the type and treatment of occurrence data used for training. The models vary in whether they incorporate urban records, apply environmental filtering, or represent a selected optimized version. Their definitions are as follows: Combined model:Modul using both natural and urban occurrence records. This model does not include any sampling bias correction. In the publication, we also refer to it as the start model. Urban model:Model using only urban occurrence records. Natural model:Model using only natural occurrence records. Filtered model:Trained with subsets of the natural occurrence records. The subsets are defined by applying an environmental filter with different partition numbers. Final model:Trained with one specific filtered subset of the natural occurrence records (based on three partitions). Data tables This section contains several tables summarizing the results and diagnostics of species distribution models (SDMs) for Aedes vexans. Below, each table is described, including the information it contains and the meaning of its column names. predictor_colinearity.csv This table shows the pairwise collinearity (Spearman correlation coefficients) between environmental predictor variables used in the models. Each row and column represents a predictor variable, and the values indicate the strength of correlation between the two variables (values close to 1 indicate high collinearity). Columns: The first column lists the predictor variable for each row. The remaining columns are predictor variables as well, matching the header row. Each cell contains the correlation coefficient between the row and column predictor. environmental_filter_models_auc_tss.csv This table summarizes the performance of different environmental filter models, reporting AUC (Area Under the Curve) and TSS (True Skill Statistic) scores for various test zones and datasets. Columns: environmental Filter parameter: The number of filter partitions. AUC - Test in Flood zones, AUC - Test in normal natural zones, AUC - Test in normal urban zones, AUC - Test both zones, AUC - Test external data: AUC scores for model performance in different test environments. TSS - Test in Flood zones, TSS - Test in normal natural zones, TSS - Test in normal urban zones, TSS - Test both zones, TSS - Test external data: TSS scores for model performance in different test environments. Integer after // indicates the nuber of model iterations in which the environmental-filtered model outperformed the model trained with all natural occurrence records for that metric and test set. combined_model_auc_results.csv This table contains the AUC score for the combined model, summarizing its overall predictive performance. Columns: AUC: The Area Under the Curve score for the combined model. combined_model_variable_importance.csv This table lists the permutation importance of each variable in the combined model, indicating how much each variable contributes to the model's predictive accuracy. The combined model was trained on the full set of natural and urban occurrence records. It represents the uncorrected baseline, as no sampling bias adjustments were applied. In the publication, we also refer to it as the start model. Columns: Variable: The name of the predictor variable. Permutation_importance: The importance score for the variable (higher means more important). sd: The standard deviation of the importance score across model runs. natural_model_variable_importance.csv This table shows the permutation importance of each variable in the natural model, indicating the contribution of each variable to the model's predictions in natural environments. Columns: Variable: The name of the predictor variable. Permutation_importance: The importance score for the variable. sd: The standard deviation of the importance score. natural_model_auc_in_flood_zones.csv This table reports the AUC scores for the natural model when tested in flood zones and natural zones across multiple iterations. Columns: iteration: The iteration number of the model run. aucFlood: The AUC score for flood zones. aucNatural: The AUC score for natural zones. The last row gives the mean and standard deviation for each zone. urban_model_variable_importance.csv This table shows the permutation importance of each variable in the urban model, indicating the contribution of each variable to the model's predictions in natural environments. Columns: Variable: The name of the predictor variable. Permutation_importance: The importance score for the variable. sd: The standard deviation of the importance score. urban_natural_model_auc_results.csv This table compares the AUC scores of natural and urban models when tested on both natural and urban datasets. Columns: Model: The type of model used (Natural or Urban). Test_Data: The type of test data (Natural or Urban). auc: The AUC score for the model on the specified test data, with mean and standard deviation. final_model_variable_importance.csv This table shows the permutation importance of each variable in the final model, indicating how much each variable contributes to the model's predictive accuracy. The final model was trained on a specific filtered subset of the natural occurrence records. Columns: Variable: The name of the predictor variable. Permutation_importance: The importance score for the variable (higher means more important). sd: The standard deviation of the importance score across model runs. predictor_metadata.ods This table provides detailed metadata for each environmental predictor variable used in the models. It describes the source, spatial and temporal resolution, extent, projection, and the processing steps applied to each variable. Columns: variableName: The name of the predictor variable. measure: Description of what the variable measures. Source: The data source (e.g., DWD, MODIS, OpenStreetMap, mundialis). rawResolution: The original spatial resolution of the data. rawExtent: The original spatial extent of the data. rawProjection: The original coordinate reference system/projection. usedResolution: The spatial resolution used in the model. usedExtent: The spatial extent used in the model. usedProjection: The coordinate reference system/projection used in the model. usedTemporalResolution: The temporal resolution used in the model (e.g., multi-annual, annual). processingSteps: Description of the processing steps applied to the raw data to produce the final predictor variable (e.g., calculation of multi-annual mean, masking, resampling, reprojection, aggregation). Prediction maps combined_model_prediction.tifRaster map showing the predicted probability of occurrence for Aedes vexans based on the combined model. Each cell value represents the model's prediction for that location. urban_model_prediction.tifRaster map showing the predicted probability of occurrence for Aedes vexans based on the urban model. Each cell value represents the model's prediction for that location. natural_model_prediction.tifRaster map showing the predicted probability of occurrence for Aedes vexans based on the natural model. Each cell value represents the model's prediction for that location. final_model_prediction.tifRaster map showing the predicted probability of occurrence for Aedes vexans based on the final selected model. Each cell value represents the model's prediction for that location. Sampling bias plots sampling_bias_before_filter.pngPlot visualizing the spatial sampling bias in the occurrence data before any filtering.Created using the sampbias R package. sampling_bias_after_filter.pngPlot visualizing the spatial sampling bias in the occurrence data after filtering.Created using the sampbias R package. sampbias R package reference: Zizka, A., Antonelli, A. and Silvestro, D. (2021), sampbias, a method for quantifying geographic sampling biases in species distribution data. Ecography, 44: 25-32. https://doi.org/10.1111/ecog.05102

提供机构:
Zenodo
创建时间:
2026-04-14
二维码
社区交流群
二维码
科研交流群
商业服务