遇见数据集

Large language models for generating disability reports and assigning impairment percentages for health commissions

收藏
Zenodo2025-05-17 更新2026-05-26 收录
官方服务:

资源简介:

This study explores the potential of Large Language Models (LLMs), specifically ChatGPT-4o and Data Analyst, in supporting health commissions with disability assessments. Prior to evaluation, interrater reliability was established. Using nine realistic patient scenarios, the artificial intelligence (AI) models were evaluated on their alignment with expert guidelines, completeness, and accuracy, using a 5-point Likert scale. The disability percentages generated by the AI systems were compared with expert-calculated values to assess their reliability. ChatGPT-4o achieved "very good" performance in most metrics, while Data Analyst was rated as "good," with no statistically significant differences observed in their overall scores. However, ChatGPT-4o failed to accurately calculate disability percentages in 5 out of 9 scenarios (55.6%), and Data Analyst failed in 8 out of 9 scenarios (88.9%). While these LLMs demonstrate the ability to produce good to very good quality reports, they currently fall short of delivering reliable disability percentage calculations, underscoring the critical need for expert supervision.

本研究探讨了大语言模型(Large Language Models, LLMs)——具体为ChatGPT-4o与Data Analyst——在协助卫生委员会开展残疾评定工作中的应用潜力。评估实施前,已预先确立评估者间信度(interrater reliability)。研究采用9个贴合临床实际的患者场景,通过5级李克特量表(Likert scale)从符合专家指南要求、内容完整性与结果准确性三个维度,对人工智能(Artificial Intelligence, AI)模型展开评估。将AI系统生成的残疾率与专家计算的基准值进行对比,以评估其可靠性。ChatGPT-4o在多数评估指标上达到了“极佳”表现,而Data Analyst获评“良好”,二者的整体得分未呈现统计学显著性差异。但在9个场景中,ChatGPT-4o有5个(占比55.6%)未能准确计算残疾率,Data Analyst则有8个(占比88.9%)出现计算失误。尽管此类大语言模型能够生成质量良好至极佳的评估报告,但目前仍无法提供可靠的残疾率计算结果,这凸显了专家监督的必要性。

提供机构:
Zenodo
创建时间:
2025-05-17
二维码
社区交流群
二维码
科研交流群
商业服务