Clinical Validity of Caries Vignettes Requires More Than Reliability
收藏资源简介:
Objectives To determine, in the same dentists, whether standardized caries cases reliably distinguish treatment tendencies, whether those tendencies recur in practice, and whether vignette-derived probabilities reproduce clinical treatment rates. Methods This same-dentist secondary analysis linked 2,207 evaluable vignette responses from 83 dentists—including a common 16-case factorial set—with lesion-level clinical records across two phases in a national practice-based network. We estimated overall and cue-specific reliability; tested recurrence in planned treatment and recorded lesion opening; and evaluated probability transfer in 1,265 intervention records from 81 dentists, excluding evaluation dentists from model fitting. Results Overall vignette treatment tendency was highly reliable (median, 0.955; dentist-cluster bootstrap 95% CI, 0.927 to 0.968), whereas cue-specific reliabilities ranged from 0.260 to 0.523. Dentists’ mean vignette treatment rates correlated with clinical treatment rates before (Spearman rho, 0.418; 95% CI, 0.233 to 0.583) and during the intervention (rho, 0.383; 95% CI, 0.184 to 0.550). From −1 to +1 SD, adjusted opening rates differed by 22 to 28 per 100 lesions. Mean held-out vignette-derived probabilities exceeded observed clinical treatment rates by 19.7 percentage points with dentist-balanced weighting and 21.6 points with case weighting and had a practitioner-balanced Brier score of 0.268; recalibration reduced the score to 0.228, closing 59.3% of the gap to the clinical logistic benchmark in the primary analysis. Clinical-record models had lower errors (0.200 to 0.201), and adding vignette records produced no demonstrated improvement. Conclusions Validity was use-specific: standardized cases reliably ranked dentists and treatment tendencies recurred in practice, but vignette probabilities overestimated clinical treatment rates. Reliable practitioner ranking and accurate clinical-rate estimation are therefore distinct validation targets for benchmarking, education, and quality improvement. Using vignette probabilities as absolute rates requires separate validation and calibration against outcomes from the intended setting.



