Dataset for Regularized Linear and Bayesian Models Improve Stacked Generalization for Species Distribution Modeling
收藏资源简介:
This deposit is the archived reproducibility package accompanying the manuscript "Dataset for Regularized Linear and Bayesian Models Improve Stacked Generalization for Species Distribution Modeling" (Payopay et al., submitted to Ecological Modelling). It contains the end-to-end species distribution modeling pipeline together with the model inputs, fitted outputs, statistical test results, and suitability maps needed to reproduce every figure and table in the paper. The study compares nine stacking meta-learners arranged along a ladder of increasing model capacity, from a regularized linear combiner through a Bayesian-linear combiner to flexible nonlinear combiners, against two traditional ensembles, namely an unweighted mean and a TSS-weighted mean. All combiners are fitted on identical out-of-fold predictions from six diverse base learners, namely Extra Trees, logistic regression, a multilayer perceptron, naive Bayes, a support vector machine, and XGBoost, so that any performance difference is attributable to the combiner alone rather than to the base ensemble. The testbed species is Bambusa oldhamii, and the study area is the province of Benguet in the northern Philippine Cordillera. The workflow incorporates spatial block cross-validation to control spatial autocorrelation, twenty pseudo-absence replicates so that pseudo-absence stochasticity propagates into all estimates, ensemble recursive feature elimination for predictor selection, a Gaussian-process meta-learner that yields an epistemic uncertainty map, and formal model comparison via the Friedman test with Nemenyi post-hoc, supplemented by Wilcoxon signed-rank and DeLong tests. The headline result is that discrimination and stability are ordered by combiner capacity rather than by algorithm family, and that a Bayesian-logistic combiner gives the best balance while capacity beyond a regularized linear combiner adds variance without improving accuracy. The archive includes the modeling script, the pinned software environment, field-collected occurrence records, and derived environmental predictor rasters. Standard public predictor layers from CHELSA, Sentinel-2, ASTER GDEM, and SoilGrids should be obtained from their original sources as cited in the manuscript, whereas the derived spectral layers are provided here. Code is released under the MIT License, and data and figures are released under the Creative Commons Attribution 4.0 International (CC BY 4.0) License. Correspondence should be directed to John Paul M. Payopay, Center for Geoinformatics, Benguet State University (jp.payopay@up.edu.ph).



