BEYOND ACCURACY
收藏资源简介:
Large language models are entering clinical practice at a pace that outstrips the development of methods for evaluating their failure modes. The principal metric by which these systems are assessed, accuracy on standardized medical question-answering benchmarks, treats all errors as equivalent, assigning the same penalty to a fabricated drug interaction as to a misidentified publication year. This equivalence is clinically untenable: the consequences of different classes of model error vary by orders of magnitude in their potential for patient harm. This paper proposes a taxonomy of five mechanistically distinct hallucination types in clinical language model output, and develops a harm-weighted scoring instrument—the Clinical Hallucination Risk Score—that evaluates model errors along three independent dimensions: clinical severity, reversibility, and propagation potential. A computational analysis applying the instrument to error distributions drawn from published clinical benchmarks demonstrates that accuracy-based rankings and harm-weighted rankings diverge substantially and systematically across hallucination types, with the models that perform best on standard benchmarks carrying disproportionately high harm-weighted risk profiles in two of the five error categories. The framework provides a methodological foundation for evaluation practices that are calibrated to the clinical consequences of model output, and generates testable predictions about the conditions under which standard evaluation is most likely to underestimate deployment risk.



