Comparative performance of ChatGPT-4.0, DeepSeek, Gemini and Perplexity in answering common questions from patients with COPD
收藏资源简介:
This repository contains the de-identified item-level data and reproducible Python code used for the revised statistical analysis of the study comparing ChatGPT-4.0, DeepSeek, Gemini, and Perplexity in answering 20 common COPD-related questions. The package includes the individual accuracy and comprehensiveness ratings provided by three COPD specialists, readability measurements, descriptive statistics, intraclass correlation coefficient analyses, Friedman tests, Holm-adjusted paired Wilcoxon signed-rank tests, exploratory domain-specific analyses, and word–character correlation analyses. The Python script reproduces the principal statistical results reported in the revised manuscript. Inter-rater reliability was evaluated using two-way random-effects, absolute-agreement intraclass correlation coefficients. The four LLMs were compared using paired analyses because the same 20 questions were evaluated across all models. No patient-level data or personally identifiable health information are included.



