A scalable Bayesian geospatial framework for small-area population estimation from multiple sparse data sources (Anonymised for Journal Reviews)
收藏资源简介:
Bayesian geospatial population modelling This repository contains the R code used for the Cameroon application, simulation study, and generation of the main population-estimation and uncertainty figures. The Cameroon analysis implements Bayesian geospatial population models using the Integrated Nested Laplace Approximation (INLA) and a Matérn spatial field represented through the SPDE approach. Two likelihood formulations are compared: (1) a Gamma model for people per building (PPB), and (2) a Negative Binomial model formulation for enumeration-area population counts with the logarithm of building count included as an offset. The candidate predictor set is restricted to eight prespecified geospatial covariates describing proximity to conflict events, explosions, water bodies, herbaceous areas, local roads and marketplaces, together with slope and night-time light intensity. Covariate selection is conducted separately for the Gamma and Negative Binomial models using a nested INLA/WAIC procedure. Importantly, selection is performed within each cross-validation training set so that held-out observations do not contribute to model selection. Predictive performance is primarily assessed using five-fold spatial block cross-validation with 100-km blocks, with random cross-validation and additional spatial sensitivity analyses used as diagnostics. The code also includes leave-region-out and leave-source-out validation, alternative spatial block sizes, prior and mesh sensitivity analyses, residual spatial-autocorrelation diagnostics, and sensitivity analyses for building counts and covariates. Following model validation, the selected Gamma and Negative Binomial models are fitted to the complete eligible dataset and used to generate gridded population predictions. For the Gamma model, predicted people-per-building values are converted to population using building counts; for the Negative Binomial model, building exposure is incorporated directly through the model offset. Cells with no mapped buildings are treated as structural zeros. Survey-source effects are excluded from national grid prediction, while uncertainty associated with the enumeration-area IID component is incorporated by marginalisation over a new zero-mean residual effect. Population uncertainty is propagated through posterior draws and aggregated to national and subnational totals. A complementary Monte Carlo simulation study evaluates the modelling framework under combinations of spatial correlation range (50, 150 and 300 km), spatial variance (0.25, 1.00 and 2.25), and enumeration-area response missingness (0%, 15% and 30%). The simulation compares the principal Gamma PPB spatial model with the Negative Binomial spatial likelihood comparator and evaluates predictive bias, MAE, RMSE, 95% interval coverage and interval width. The accompanying plotting code reproduces comparisons of Gamma and Negative Binomial population estimates, credible intervals, coefficients of variation, regional and national estimates, and spatial population and uncertainty maps. The deposited scripts use relative file paths rather than the original project-specific storage locations. Input datasets that cannot be redistributed must therefore be placed in the corresponding local data directories before execution.



