Large language models for generating disability reports and assigning impairment percentages for health commissions
收藏资源简介:
This study explores the potential of Large Language Models (LLMs), specifically ChatGPT-4o and Data Analyst, in supporting health commissions with disability assessments. Prior to evaluation, interrater reliability was established. Using nine realistic patient scenarios, the artificial intelligence (AI) models were evaluated on their alignment with expert guidelines, completeness, and accuracy, using a 5-point Likert scale. The disability percentages generated by the AI systems were compared with expert-calculated values to assess their reliability. ChatGPT-4o achieved "very good" performance in most metrics, while Data Analyst was rated as "good," with no statistically significant differences observed in their overall scores. However, ChatGPT-4o failed to accurately calculate disability percentages in 5 out of 9 scenarios (55.6%), and Data Analyst failed in 8 out of 9 scenarios (88.9%). While these LLMs demonstrate the ability to produce good to very good quality reports, they currently fall short of delivering reliable disability percentage calculations, underscoring the critical need for expert supervision.



