遇见数据集

Large language models for generating disability reports and assigning impairment percentages for health commissions

收藏
Zenodo2025-05-17 更新2026-05-26 收录
官方服务:

资源简介:

This study explores the potential of Large Language Models (LLMs), specifically ChatGPT-4o and Data Analyst, in supporting health commissions with disability assessments. Prior to evaluation, interrater reliability was established. Using nine realistic patient scenarios, the artificial intelligence (AI) models were evaluated on their alignment with expert guidelines, completeness, and accuracy, using a 5-point Likert scale. The disability percentages generated by the AI systems were compared with expert-calculated values to assess their reliability. ChatGPT-4o achieved "very good" performance in most metrics, while Data Analyst was rated as "good," with no statistically significant differences observed in their overall scores. However, ChatGPT-4o failed to accurately calculate disability percentages in 5 out of 9 scenarios (55.6%), and Data Analyst failed in 8 out of 9 scenarios (88.9%). While these LLMs demonstrate the ability to produce good to very good quality reports, they currently fall short of delivering reliable disability percentage calculations, underscoring the critical need for expert supervision.

提供机构:
Zenodo
创建时间:
2025-05-17
二维码
社区交流群
二维码
科研交流群
商业服务