HealthChat-11K
收藏资源简介:
HealthChat-11K 是一个由 UNC Chapel Hill, Duke University, University of Washington 和 Google 研究人员合作创建的,包含 11,000 个真实对话,共计 25,000 条用户消息的数据集。该数据集旨在研究用户在与大型语言模型进行健康信息交流时的互动模式,并通过对 21 个不同健康专业的用户互动进行分类和注释,以帮助理解和改进聊天机器人在医疗保健支持方面的能力。数据集的创建过程包括从大规模对话数据集中筛选健康相关对话,并使用 LLM 进行对话级别的专业分类和消息级别的分类代码注释。该数据集可用于研究用户如何与聊天机器人进行多轮对话,以及如何提出开放式、常常模糊的问题。数据集的最终目标是促进对大型语言模型在医疗保健领域应用的更深入理解,并帮助改善其支持能力。
HealthChat-11K is a dataset co-created by researchers from UNC Chapel Hill, Duke University, the University of Washington, and Google, containing 11,000 real conversations with a total of 25,000 user messages. This dataset aims to study the interaction patterns of users when exchanging health information with large language models, and to help understand and improve the capabilities of chatbots in healthcare support by categorizing and annotating user interactions across 21 different health specialties. The dataset creation process involves screening health-related conversations from large-scale conversational datasets, and using LLMs to perform specialty classification at the conversation level and annotate classification codes at the message level. This dataset can be used to study how users conduct multi-turn conversations with chatbots and how they raise open-ended, often ambiguous questions. The ultimate goal of this dataset is to promote a deeper understanding of the applications of large language models in the healthcare field and help improve their support capabilities.
HealthChat数据集概述
数据集基本信息
- 项目名称:HealthChat
- 发布状态:持续更新中(截至2025年7月1日发布HealthChat-11K)
- 主要目标:改善人类与AI(如大型语言模型)之间的健康对话
数据集内容
- 数据规模:11,000条健康对话记录(HealthChat-11K)
- 数据特点:真实场景下的用户健康对话交互数据
相关研究支持
- 配套文献:配套arXiv论文(未提供具体论文标题)
- 研究用途:支持健康对话领域的用户交互研究
数据获取
- 发布渠道:通过GitHub项目发布(https://github.com/yahskapar/HealthChat)

- 1"What's Up, Doc?": Analyzing How Users Seek Health Information in Large-Scale Conversational AI DatasetsUNC Chapel Hill, Duke University, University of Washington, Google · 2025年



