A Reproducible Bayesian–Mechanistic Framework for Reproduction-Number Estimation and Early Warning in Resource-Limited Surveillance: An Illustrative Simulation Study
收藏资源简介:
Advanced epidemic-modelling methods are often difficult to deploy in resource-limited settings, where surveillance is weekly, delayed, and substantially under-reported. We present the Hybrid Bayesian–Mechanistic Framework (HBMF), which couples a mechanistic Erlang–SEIR model with a penalised B-spline (P-spline) representation of the time-varying transmission rate, a reporting-delay and time-varying reporting observation model, and a negative-binomial likelihood, and which provides fully-marginalised uncertainty through Markov chain Monte Carlo (MCMC). The core of this paper is a declared, fully synthetic identifiability and calibration study: every synthetic-study figure and quantity is produced by the accompanying open-source code from data generated under a known ground truth, with no external benchmark comparison and no headline predictive-accuracy percentage. On synthetic weekly, delayed, under-reported data (mean reporting fraction 0.42), HBMF recovers the reproduction-number trajectory (posterior median RMSE 0.15; on the reference dataset the 95% credible band contains the true R_t at 100% of days) and, across 44 independent replicate datasets, yields well-calibrated posterior-predictive intervals for reported cases (empirical coverage 0.52/0.82/0.97 at nominal 0.50/0.80/0.95). We are equally explicit about the method's limits: the level of R_t and the reporting trend are only partially identified from a single incidence series, and a fast Laplace approximation under-covers latent R_t (its 95% intervals are 0.58× the width of the full-MCMC intervals). Convergence is verified (maximum split-R̂=1.06; minimum effective sample size 166). Variance-based global sensitivity attributes 91%–99% of peak-incidence and final-size variance to R_0. A model-based early-warning score discriminates large from limited outbreaks using only five weeks of data with area under the ROC curve 0.90. We additionally report one supplementary, single-city illustrative application to genuine weekly dengue surveillance data (Rio de Janeiro, 2021; InfoDengue), explicitly not presented as a validation, since no true R_t exists for real data: in-sample posterior-predictive coverage was good and Bayesian p-values showed no gross misfit, and agreement with an independent operational R_t estimate was moderate (r=0.43). An initial random-walk-Metropolis fit converged markedly less cleanly on this real series than on synthetic data; implementing a Laplace-preconditioned Hamiltonian Monte Carlo sampler in the same validated codebase reduced the maximum split-R̂ from 1.57 to 1.18 using under 7% of the MCMC draws. A genuine out-of-sample forecast test on six held-out weeks was outperformed by a naive baseline, a result traceable to the same boundary-identifiability weakness documented in the synthetic study; using the properly converged posterior for this test revealed that the initial, less-converged chain had understated the model's true forecast uncertainty by roughly an order of magnitude, which we regard as the paper's clearest demonstration that convergence diagnostics carry real operational consequences. All code, the exact real-data retrieval query, and the HMC implementation are released to enable exact reproduction.



