遇见数据集

Accuracy of ChatGPT in answering cardiology board-style questions

收藏
DataONE2025-02-26 更新2025-12-06 收录
官方服务:

资源简介:

Studies were included in the analysis if they met all of the following criteria: (1) the article was written in English, (2) the study assessed ChatGPT’s accuracy on questions that were set at a level or retrieved from an appropriate resource representing board-style (specialist) cardiology certification examination questions, (3) the questions inputted into ChatGPT were either text-based, image-based, or a combination of both, and (4) the study provided data on the number of questions inputted into ChatGPT and the number (or percentage) of correct responses reported separately for each question format. Studies were excluded if they failed to meet any of the aforementioned inclusion criteria or did not disclose original data, such as review papers or descriptive replies/correspondence to previously published articles. Key study characteristic data from included studies were extracted and entered into a predefined data abstraction template. Statistical analysis pooled the reported accuracy from each study to calculate an overall pooled accuracy with a 95% confidence interval (CI), subgrouped by model version, using a random-effects model. The meta-analysis software used was STATA ver. 18.0 (Stata Corp.). P-values <0.05 were considered statistically significant. Heterogeneity was assessed using the I2 statistic.

本分析纳入符合以下全部标准的研究:(1) 研究论文以英文撰写;(2) 研究评估了ChatGPT在匹配特定难度层级,或源自恰当资源的模拟专科心血管病学委员会认证考试风格试题上的准确率;(3) 输入ChatGPT的试题可采用文本形式、图像形式,或二者结合的形式;(4) 研究需提供输入ChatGPT的试题总数量,以及按每种试题格式分别统计的正确作答数量(或占比)数据。 凡未满足上述任一纳入标准,或未披露原始数据的研究(如综述类论文、针对已发表文章的评述/通信类文章),均予以排除。 从纳入的研究中提取关键研究特征数据,并录入预设的数据提取模板。采用随机效应模型,按模型版本进行亚组分析,对各研究报告的准确率进行合并计算,以得出合并后的总准确率及其95%置信区间(CI)。本研究使用的元分析软件为STATA 18.0版(Stata公司)。将P值<0.05判定为具有统计学显著性。异质性检验采用I²统计量。

创建时间:
2025-10-29
二维码
社区交流群
二维码
科研交流群
商业服务