ThReadMed-QA
收藏资源简介:
ThReadMed-QA是由东北大学·科里计算机科学学院创建的多轮医疗对话数据集,旨在评估大语言模型在真实患者互动中识别和纠正误解的能力。该数据集包含2,437个对话线程和8,204个问答对,数据来源于Reddit的r/AskDocs子论坛,覆盖了2015年至2023年间的患者与医生自然交互,平均患者提问长度为129个tokens,医生回复为75个tokens。其构建过程通过筛选原始帖子、预处理、提取患者与医生对话分支,并过滤低质量内容,最终保留完全回答的线程。该数据集主要应用于医疗人工智能领域,旨在解决大语言模型在连续对话中处理患者误解的可靠性问题,为评估模型在多轮互动中的安全性和一致性提供基准。
ThReadMed-QA is a multi-turn medical dialogue dataset created by Khoury College of Computer Sciences, Northeastern University, aiming to evaluate the ability of Large Language Models (LLMs) to recognize and correct misunderstandings during real patient interactions. This dataset contains 2,437 dialogue threads and 8,204 question-answer pairs, sourced from the r/AskDocs subreddit of Reddit, covering natural patient-doctor interactions between 2015 and 2023, with an average length of 129 tokens for patient queries and 75 tokens for physician responses. Its construction pipeline includes screening original posts, preprocessing, extracting patient-doctor dialogue branches, filtering low-quality content, and ultimately retaining fully answered dialogue threads. This dataset is primarily applied in the field of medical artificial intelligence, aiming to address the reliability issues of LLMs when handling patient misunderstandings in continuous conversations, and providing a benchmark for evaluating the safety and consistency of models in multi-turn interactions.

- 1Evaluating Large Language Models on Misconceptions in Multi-Turn Medical Conversations东北大学·科里计算机科学学院 · 2026年




