AI vs. Human Experts Critical Appraisal Fresno Scores
收藏资源简介:
Objective We aimed to test the hypothesis that a significant difference exists between artificial intelligence and human experts in conducting critical appraisal of standard study designs as assessed by questions 6 and 7 of the validated FRESNO rubric. Methods Each human expert conducted a critical appraisal of three pre-selected studies: one RCT, one systematic review, and one diagnosis study. They then prompted their assigned AI model to perform the same task using a structured prompt. Participants submitted all critical appraisal artifacts to the PI, who de-identified each artifact. The artifacts were assigned to a separate and independent group of graders, who were trained and normed in the application of Fresno questions 6 and 7 to grade each artifact. Once all data was collected and scored by trained raters, the data analyst processed the information for group differences using Welch's ANOVA. Results After the data analysis, the team confirmed that not only was the AI models' performance at the critical appraisal task superior to the human experts across the board, but that the AI models also performed the task with less variance in performance as well. On a 36-point scale, the AI models scored a mean of 31.4 with a standard deviation of 7.1, whereas the human experts scored a mean of 21.3 with a standard deviation of 11.2. Between groups we observed a statistically significant mean score difference of M=10.1 (95% CI [5.4 to 14.7], p<.001). Of the models, ChatGPT scored the highest with a mean of 32.92, followed by Copilot with a mean of 30.67, and the lowest performer was Gemini with a mean of 30.5. However, all these models’ mean scores were still higher than the human experts’ mean. Conclusion Overall, AI-generated critical appraisals can be helpful in adding important context to the interpretation of research studies, especially when aiming to evaluate methodological rigor and magnitude of effect size. As AI summaries and clinical decision-making tools become more prevalent, understanding the quality and rigor of AI output compared to human expertise will become increasingly important.



