Study characteristics.
收藏资源简介:
Despite advances in deep learning and transformer architectures, prior reviews have focused narrowly on traditional clinical decision support systems (CDSS) or single medical domains, leaving significant gaps in understanding contemporary AI-driven predictive tools. This systematic review and meta-analysis evaluated the predictive performance of artificial intelligence-based CDSS (AI-CDSS) across multiple medical specialties. Following PRISMA guidelines, PubMed and Cochrane Library were searched through December 2024 for studies evaluating predictive AI-CDSS using real-world clinical data. Two reviewers independently screened 3,296 records (κ = 0.833), with study quality assessed via QUADAS-2 and performance measures pooled using random-effects meta-analysis. Fifty studies spanning 17 medical specialties were included. Meta-analysis demonstrated moderate discriminatory ability (pooled AUC: 0.652, 95% CI: 0.562–0.743), high specificity (0.819, 95% CI: 0.793–0.844), moderate accuracy (0.765, 95% CI: 0.734–0.796), and variable sensitivity (0.660, 95% CI: 0.535–0.785), with substantial heterogeneity across all measures (I² ≥ 98.9%). Only 24% of studies involved prospective deployment, and 64% reported exclusively technical metrics without clinical workflow data. Predictive AI-CDSS demonstrate moderate-to-good diagnostic performance with strong specificity; however, the predominance of retrospective study designs and limited implementation reporting reveal critical gaps between technical validation and real-world clinical utility. To address these shortcomings, we propose the ROADMAP framework, structured around seven domains: Representative development, Outcomes-focused evaluation, Assessment for deployment, Data harmonization, Monitoring for bias, Allocation via economic evaluations, and Priorities for standardized reporting and prospective validation. This framework provides a practical roadmap for bridging the gap between algorithmic performance and meaningful clinical integration.
尽管深度学习与Transformer架构已取得诸多进展,但既往综述往往仅聚焦于传统临床决策支持系统(clinical decision support system, CDSS)或单一医学领域,导致学界对当代人工智能驱动的预测工具的认知存在显著缺口。本项系统综述与荟萃分析针对跨多医学专科的人工智能辅助临床决策支持系统(AI-CDSS)的预测性能展开评估。本研究遵循PRISMA指南,对PubMed及Cochrane Library数据库截至2024年12月收录的、采用真实世界临床数据评估预测型AI-CDSS的相关研究进行了检索。两名研究者独立筛选了3296条记录(κ=0.833),采用QUADAS-2工具评估研究质量,并通过随机效应荟萃分析合并性能指标。最终共纳入覆盖17个医学专科的50项研究。荟萃分析结果显示,该类系统展现出中等程度的区分能力(合并AUC:0.652,95%CI:0.562~0.743)、较高的特异性(0.819,95%CI:0.793~0.844)、中等的准确率(0.765,95%CI:0.734~0.796),以及波动范围较大的灵敏度(0.660,95%CI:0.535~0.785);所有指标均存在显著异质性(I²≥98.9%)。仅24%的研究涉及前瞻性部署,且64%的研究仅报告了技术指标,未提供临床工作流相关数据。预测型AI-CDSS展现出中等至良好的诊断性能,且特异性较强;但现有研究多为回顾性设计,且对系统落地情况的报告较为有限,这暴露出算法技术验证与真实世界临床应用之间存在的关键鸿沟。为弥补上述不足,本研究提出ROADMAP框架,该框架围绕七大维度构建:代表性开发、结局导向型评估、部署前评估、数据协调、偏倚监测、基于经济学评价的资源分配,以及标准化报告与前瞻性验证的优先级。该框架可为弥合算法性能与临床实际整合之间的差距提供切实可行的路线图。



