遇见数据集

A Trust-Calibrated, Causally-Discounted Framework for Cross-Scale Evidence Fusion Under Data Bias, Population Heterogeneity, and Temporal Drift

收藏
Zenodo2026-07-26 更新2026-08-01 收录
官方服务:

资源简介:

Clinical and biomedical inference increasingly depends on evidence collected across heterogeneous scales -- genomic, proteomic, cellular, organ-level, and whole-patient -- yet no unified formalism exists for combining such evidence when it is simultaneously threatened by (i) fixed measurement and selection bias, (ii) population heterogeneity between the source and target cohorts, (iii) temporal (data) drift, and (iv) uncertain causal identifiability, i.e., the persistent ambiguity between correlational and causal support for a given source of evidence. A fifth, rarely modeled factor compounds all four: the treating physician's trust in an aggregated recommendation does not track system accuracy instantaneously, but is itself a dynamical, hysteretic process that can amplify or dampen the practical consequences of evidentiary error. This paper proposes the Multi-Scale Evidentiary Coherence with Trust Calibration (MSEC-TC) framework, which (a) formalizes a validity-discounted, precision-weighted fusion rule -- the Cross-Scale Bayesian Discounted Fusion (CSBDF) estimator -- that jointly discounts each evidentiary source by its causal identifiability, standardized population shift, and a population-stability-index-based drift statistic; (b) introduces the Cross-Scale Causal Coherence Index (C3I), a real-time, interpretable diagnostic of how much of a fused estimate's precision is contributed by causally valid, non-drifted sources; and (c) couples the fusion layer to a delay-hysteresis ordinary differential equation governing physician trust, closing the evidence-to-decision loop. We derive a bias bound (Theorem 1) showing that CSBDF's worst-case fused error is provably no larger, and generically strictly smaller, than that of naive inverse-variance-weighted fusion whenever causal validity is heterogeneous across scales. We instantiate the framework in a fully reproducible synthetic five-scale system (gene, protein, cell, organ, patient) and benchmark CSBDF against three literature-representative baselines -- naive inverse-variance-weighted fusion, fixed equal-weight ensembling, and importance-weighted covariate-shift correction -- across three qualitatively distinct regimes. CSBDF reduces expected calibration error relative to naive fusion by 48.5% in a confounding-dominant regime and by 32.9% in a mixed regime, and a dedicated heterogeneity sweep confirms the theoretically predicted monotonic relationship between causal-validity heterogeneity and CSBDF's advantage (a falsifiable claim we test directly). We report, with equal candor, that CSBDF does not uniformly dominate on raw point-error (Brier score) -- a simpler shift-only correction is competitive or superior in low-confounding, drift-dominant regimes, and a 3×3×3 robustness grid shows CSBDF is best-or-tied on Brier score in only 25.9% of parameter combinations, while it improves calibration in 82.6% of 500 Monte Carlo replications. Multi-parameter Sobol sensitivity analysis identifies confounder loading and fixed bias magnitude, not drift rate, as the dominant drivers of CSBDF's advantage (total-order indices 0.462 and 0.484, respectively). We conclude with an explicit falsifiability program, a scientific and technical risk assessment, and a staged roadmap for prospective experimental validation, positioning MSEC-TC as a conceptual, hypothesis-generating contribution rather than a deployment-ready clinical tool.

提供机构:
Zenodo
创建时间:
2026-07-26
二维码
社区交流群
二维码
科研交流群
商业服务