遇见数据集

A comparative evaluation of ChatGPT 3.5 and ChatGPT 4 in responses to selected genetics questions - Full study data

收藏
DataONE2024-06-04 更新2025-08-02 收录
官方服务:

资源简介:

Objective: Our objective is to evaluate the efficacy of ChatGPT 4 in accurately and effectively delivering genetic information, building on previous findings with ChatGPT 3.5. We focus on assessing the utility, limitations, and ethical implications of using ChatGPT in medical settings. Materials and Methods: A structured questionnaire, including the Brief User Survey (BUS-15) and custom questions, was developed to assess ChatGPT 4's clinical value. An expert panel of genetic counselors and clinical geneticists independently evaluated ChatGPT 4's responses to these questions. We also involved comparative analysis with ChatGPT 3.5, utilizing descriptive statistics and using R for data analysis. Results: ChatGPT 4 demonstrated improvements over 3.5 in context recognition, relevance, and informativeness. However, performance variability and concerns about the naturalness of the output were noted. No significant difference in accuracy was found between ChatGPT 3.5 and 4.0. Notably, the effic..., Study Design This study was conducted to evaluate the performance of ChatGPT 4 (March 23rd, 2023)  Model) in the context of genetic counseling and education. The evaluation involved a structured questionnaire, which included questions selected from the Brief User Survey (BUS-15) and additional custom questions designed to assess the clinical value of ChatGPT 4's responses. Questionnaire Development The questionnaire was built on Qualtrics, which comprised twelve questions: seven selected from the BUS-15 preceded by two additional questions that we designed. The initial questions focused on quality and answer relevancy: 1.    The overall quality of the Chatbot’s response is: (5-point Likert: Very poor to Very Good) 2.    The Chatbot delivered an answer that provided the relevant information you would include if asked the question. (5-point Likert: Strongly disagree to Strongly agree) The BUS-15 questions (7-point Likert: Strongly disagree to Strongly agree) focused on: 1.    Recogniti..., , # A comparative evaluation of ChatGPT 3.5 and ChatGPT 4 in responses to selected genetics questions - Full study data [https://doi.org/10.5061/dryad.s4mw6m9cv](https://doi.org/10.5061/dryad.s4mw6m9cv) This data was captured when evaluating the ability of ChatGPT to address questions patients may ask it about three genetic conditions (BRCA1, HFE, and MLH1). This data is associated with the JAMIA article of the similar name with the DOI 10.1093/jamia/ocae128 ## Description of the data and file structure 1. **Key**: This tab contains the data structure, explaining the survey questions, and potential responses available. 2. **Prompt Responses**: This tab contains the prompts used for ChatGPT, and the response provided from each model (3.5 and 4) 3. **GPT 4 Results**: This tab provides the responses collected from the medical experts (genetic counselors and clinical geneticist) from the Qualtrics survey. 4. **Accuracy (Qx_1)**: This tab contains the subset of results from both the Ch...

研究目标:本研究基于ChatGPT 3.5的既往研究结果,评估ChatGPT 4准确且高效传递遗传信息的效能。本研究重点评估ChatGPT在医疗场景中应用的实用性、局限性及伦理影响。 材料与方法:本研究开发了包含简版用户调查(Brief User Survey, BUS-15)与自定义问题的结构化问卷,用于评估ChatGPT 4的临床应用价值。由遗传咨询师与临床遗传学家组成的专家小组,独立评估ChatGPT 4对上述问卷问题的应答情况。本研究同时开展与ChatGPT 3.5的对比分析,采用描述性统计方法,并使用R语言进行数据分析。 研究结果:ChatGPT 4在上下文识别、应答相关性与信息丰富度方面均优于ChatGPT 3.5。但研究同时观察到模型性能存在波动,且存在输出自然度相关的顾虑。ChatGPT 3.5与4.0的应答准确性无显著差异。值得注意的是,本研究效能[原文截断]。 研究设计:本研究旨在评估2023年3月23日发布的ChatGPT 4模型在遗传咨询与科普场景中的表现。评估采用结构化问卷,问卷包含从简版用户调查(BUS-15)中选取的题目,以及本研究设计的自定义题目,用于评估ChatGPT 4应答的临床价值。 问卷编制:本研究基于Qualtrics平台编制问卷,共包含12道题目:2道自定义前置题目,以及从简版用户调查(BUS-15)中选取的7道题目。前置题目聚焦应答质量与相关性: 1. 该聊天机器人应答的整体质量为:(5级李克特量表:极差至极佳) 2. 该聊天机器人所提供的应答包含了提问时所需的相关信息。(5级李克特量表:完全不同意至完全同意) 简版用户调查(BUS-15)的题目采用7级李克特量表(完全不同意至完全同意),聚焦于:1. 识别[原文截断] # 针对遗传学相关问题的ChatGPT 3.5与ChatGPT 4应答对比评估——完整研究数据集 本数据集关联的学术文章DOI为:https://doi.org/10.5061/dryad.s4mw6m9cv。本数据采集自针对ChatGPT解答患者关于三种遗传疾病(BRCA1、HFE及MLH1)相关疑问的能力评估过程。本数据集与同名期刊文章(发表于JAMIA,DOI:10.1093/jamia/ocae128)相关联。 ## 数据与文件结构说明 1. **数据说明页(Key)**:本页包含数据结构说明,对问卷题目及可选应答选项进行解释。 2. **提示词与应答(Prompt Responses)**:本页包含用于ChatGPT的提示词,以及两款模型(3.5与4)生成的应答内容。 3. **GPT 4 评估结果(GPT 4 Results)**:本页包含医学专家(遗传咨询师与临床遗传学家)通过Qualtrics问卷收集的评估结果。 4. **准确性(Qx_1)**:本页包含两款模型的部分评估结果[原文截断]

创建时间:
2025-08-01
二维码
社区交流群
二维码
科研交流群
商业服务