Queries entered into GenAI tools.
收藏资源简介:
Generative artificial intelligence (genAI) has potential to improve healthcare by reducing clinician burden and expanding services, among other uses. There is a significant gap between the need for mental health care and available clinicians in the United States–this makes it an attractive target for improved efficiency through genAI. Among the most sensitive mental health topics is suicide, and demand for crisis intervention has grown in recent years. We aimed to evaluate the quality of genAI tool responses to suicide-related queries. We entered 10 suicide-related queries into five genAI tools–ChatGPT 3.5, GPT-4, a version of GPT-4 safe for protected health information, Gemini, and Bing Copilot. The response to each query was coded on seven metrics including presence of a suicide hotline number, content related to evidence-based suicide interventions, supportive content, harmful content. Pooling across tools, most of the responses (79%) were supportive. Only 24% of responses included a crisis hotline number and only 4% included content consistent with evidence-based suicide prevention interventions. Harmful content was rare (5%); all such instances were delivered by Bing Copilot. Our results suggest that genAI developers have taken a very conservative approach to suicide-related content and constrained their models’ responses to suggest support-seeking, but little else. Finding balance between providing much needed evidence-based mental health information without introducing excessive risk is within the capabilities of genAI developers. At this nascent stage of integrating genAI tools into healthcare systems, ensuring mental health parity should be the goal of genAI developers and healthcare organizations.
生成式人工智能(Generative Artificial Intelligence,genAI)具备通过减轻临床医师工作负担、拓展服务范畴等多种途径改善医疗健康领域的潜力。当前美国精神卫生服务需求与可用临床医师数量之间存在显著缺口,这使得genAI成为提升诊疗效率的极具吸引力的应用方向。自杀属于敏感度极高的精神卫生议题之一,近年来危机干预的需求持续增长。本研究旨在评估genAI工具针对自杀相关咨询的回复质量。我们将10个自杀相关咨询问题输入至5款genAI工具中,分别为ChatGPT 3.5、GPT-4、适配受保护健康信息(Protected Health Information)的GPT-4安全版本、Gemini以及必应Copilot(Bing Copilot)。我们针对每个回复的7项指标进行编码,涵盖是否包含自杀求助热线号码、是否涉及基于证据的自杀干预相关内容、是否为支持性内容以及是否存在有害内容等维度。综合所有工具的回复来看,绝大多数回复(79%)属于支持性内容。仅24%的回复包含危机求助热线号码,仅有4%的回复内容符合基于证据的自杀预防干预规范。有害内容极为罕见(占比5%),且全部由必应Copilot生成。本研究结果显示,genAI开发者针对自杀相关内容采取了极为保守的策略,仅限制模型回复以引导用户寻求帮助,其余相关信息则极少提供。在提供亟需的基于证据的精神卫生信息与避免过度风险之间找到平衡点,是genAI开发者能够实现的目标。在将genAI工具整合至医疗系统的初期阶段,保障精神卫生服务的公平性应当成为genAI开发者与医疗保健机构的共同目标。




