A drift-test audit for frozen genomic variant-effect classifiers: application to AlphaMissense against nine years of ClinVar reclassification
收藏资源简介:
A variant-effect classifier is typically frozen at release, while the clinical adjudications used to judge it keep moving as new evidence accrues. We study a directly matched genomic monitoring problem using standard sequential-testing tools and apply it to a decade-scale natural experiment: AlphaMissense pathogenicity scores from the released score table, thresholded by a fixed numeric convention, tracked against ClinVar adjudications recorded in 108 monthly snapshots from January 2016 to December 2024 (the complete set of monthly NCBI ClinVar variant-summary archive snapshots available over this period). Restricting to the n=7,132 SNVs that already carried an AlphaMissense score and were non-definitive at a 2016 ClinVar-label freeze, then later reached a definitive ClinVar call (pathogenic/likely pathogenic or benign/likely benign), the frozen classifier's error rate against later-adjudicated labels is 14.68% (1,047/7,132; Wilson 95% CI [13.88%,15.52%]), with errors skewed toward one failure direction (326 false-pathogenic-direction versus 721 false-benign-direction). A planned Waudby–Smith–Ramdas bounded-mean confidence-sequence component failed on this cohort: the discretized-grid implementation's surviving candidate-mean set is exhausted before the final accrual point (empty by t=1,447 of 7,132), so no bound or estimate from that component is claimed here; we flag this as an implementation defect uncovered by the real run rather than report a substitute number. A betting-martingale drift test, run against three later split points (December 2019, December 2020, December 2021) plus an early-window split, finds no evidence of drift at any of the four split points tested: the peak log-capital across all four splits is 0.81 (early-window) / 0.63 (later splits, maximum), against the per-split Ville threshold (1/)=3.00 at =0.05; unlike an earlier smaller-cohort run, the early-window split does not cross the threshold here. A same-data permutation calibration self-check (2000 trials) gives a false-alarm rate of 6.85% against a nominal =0.05, so it is reported as a diagnostic warning rather than as proof of exact calibration. A minimum-detectable-effect analysis on this cohort shows the drift test reaches 80% power at a post-split error-rate multiplier of roughly 1.23–1.26 (a relative rise from 14.7% to about 18%), so the non-detection is a bounded negative rather than a vacuously underpowered one. The result is a negative drift-test audit finding on the component that could be validly computed: a widely used frozen pathogenicity score table shows no detected error-rate drift against an evolving expert-curated label stream over nine years of monthly accrual. The failed component requires replacement by a numerically stable, independently validated implementation before it can support any time-uniform error-rate statement. The contribution is not the underlying sequential-testing machinery, which is standard, but the monitoring object itself: a drift-test audit for frozen genomic variant-effect classifiers with explicit predictor/label-stream non-feedback separation, together with an apparently first genomic application found in our literature search and a report of both the no-drift finding and an implementation failure mode encountered at full cohort scale.



