A Reproducible Bayesian–Mechanistic Framework for Reproduction-Number Estimation and Early Warning in Resource-Limited Surveillance: An Illustrative Simulation Study
收藏资源简介:
Advanced epidemic-modelling methods are often difficult to deploy in resource-limited settings, where surveillance is weekly, delayed, and substantially under-reported. We present the Hybrid Bayesian–Mechanistic Framework (HBMF), which couples a mechanistic Erlang–SEIR model with a penalised B-spline (P-spline) representation of the time-varying transmission rate, a reporting-delay and time-varying reporting observation model, and a negative-binomial likelihood, and which provides fully-marginalised uncertainty through Markov chain Monte Carlo (MCMC). The core of this paper is a declared, fully synthetic identifiability and calibration study: every synthetic-study figure and quantity is produced by the accompanying open-source code from data generated under a known ground truth; no headline predictive-accuracy percentage is reported. On synthetic weekly, delayed, under-reported data (mean reporting fraction 0.42), HBMF recovers the reproduction-number trajectory (posterior median RMSE 0.15; on the reference dataset the 95% credible band contains the true R_t at 100% of days) and, across 44 independent replicate datasets, yields well-calibrated posterior-predictive intervals for reported cases (empirical coverage 0.52/0.82/0.97 at nominal 0.50/0.80/0.95). The method's limits are equally clear: the level of R_t and the reporting trend are only partially identified from a single incidence series, and a fast Laplace approximation under-covers latent R_t (its 95% intervals are 0.58× the width of the full-MCMC intervals). Convergence is verified (maximum split-R̂=1.06; minimum effective sample size 166). Variance-based global sensitivity attributes 91%–99% of peak-incidence and final-size variance to R_0. A model-based early-warning score discriminates large from limited outbreaks using only five weeks of data with area under the ROC curve 0.90. In a single external comparison, against EpiEstim given its most favourable possible treatment (the true generation-time distribution) on the same weekly, delayed, under-reported data, pooled over the same 44 replicates HBMF achieves both lower error (0.149 vs. EpiEstim's best-case 0.338) and far better-calibrated 95% intervals (0.813 vs. 0.069 coverage); EpiEstim's coverage degrades further as its smoothing window widens, consistent with its implicit assumption that R_t is constant within each window. One supplementary, single-city illustrative application to real weekly dengue surveillance data (Rio de Janeiro, 2021; InfoDengue) is explicitly not presented as a validation, since no true R_t exists for real data: in-sample posterior-predictive coverage was good and Bayesian p-values showed no gross misfit, and agreement with an independent operational R_t estimate was moderate (r=0.43). An initial random-walk-Metropolis fit converged markedly less cleanly on this real series than on synthetic data; implementing a Laplace-preconditioned Hamiltonian Monte Carlo sampler in the same validated codebase reduced the maximum split-R̂ from 1.57 to 1.18 using under 7% of the MCMC draws. An out-of-sample forecast test on six held-out weeks was outperformed by a naive baseline, a result traceable to the same boundary-identifiability weakness documented in the synthetic study; using the properly converged posterior for this test revealed that the initial, less-converged chain had understated the model's true forecast uncertainty by roughly an order of magnitude, demonstrating that convergence diagnostics carry real operational consequences. All code, the exact real-data retrieval query, and the HMC implementation are released to enable exact reproduction.



