遇见数据集

MedDialog

收藏
魔搭社区2026-07-13 更新2025-03-01 收录
官方服务:

资源简介:

# MedDialog A large-scale medical dialogue dataset containing ~252k patient-doctor conversation pairs for training and evaluating clinical dialogue systems. ## Dataset Description | Property | Value | |---|---| | **Source** | [ruslanmv/ai-medical-chatbot](https://huggingface.co/datasets/ruslanmv/ai-medical-chatbot) | | **License** | [Apache-2.0](https://www.apache.org/licenses/LICENSE-2.0) | | **Language** | English | | **Total examples** | 251,731 | | **Train split** | 226,557 | | **Validation split** | 25,174 | | **Domain** | Clinical / General Medicine | ### Overview MedDialog is designed for training language models to generate medically appropriate, empathetic, and safe doctor responses given a patient query. It covers a broad range of medical topics including symptoms, diagnoses, medications, procedures, and general health questions. ## Dataset Structure ### Fields | Field | Type | Description | |---|---|---| | `patient_message` | `string` | The patient's description of symptoms or medical question | | `doctor_response` | `string` | The doctor's response (ground truth reference) | | `dialogue_context` | `string` | Additional context (reserved for multi-turn; empty in current version) | ### Statistics | Metric | Patient Message | Doctor Response | |---|---|---| | Min length (chars) | 22 | 41 | | Median length (chars) | 414 | 481 | | Mean length (chars) | 492 | 546 | | Max length (chars) | 2,500 | 2,500 | ## Example **Patient:** > I have been experiencing severe headaches for the past week, mostly in the morning. The pain is throbbing and located on the right side of my head. I also feel nauseous sometimes. What could be causing this? **Doctor:** > I understand how concerning persistent headaches can be, especially when accompanied by nausea. Based on your description of throbbing, one-sided pain with morning occurrence, this could potentially be migraines, though we'd need a proper evaluation to be sure. > > I'd recommend: > 1. Keep a headache diary noting triggers, duration, and severity > 2. Ensure you're staying hydrated and getting adequate sleep > 3. Avoid known migraine triggers like bright lights or certain foods > > However, given the duration and severity, I strongly advise scheduling an appointment with your doctor for a proper examination. ## Data Processing This dataset was derived from [ruslanmv/ai-medical-chatbot](https://huggingface.co/datasets/ruslanmv/ai-medical-chatbot) (257k raw examples) with the following processing steps: 1. **Field combination**: Merged `Description` and `Patient` fields into `patient_message` 2. **Quality filtering**: Removed examples with very short messages (<5 words patient, <10 words doctor) 3. **Redirect filtering**: Excluded entries where the doctor response was only a referral with no content 4. **Truncation**: Capped messages at 2,500 characters 5. **Split**: 90/10 train/validation split with random seed 42 ## Usage ### Loading with `datasets` ```python from datasets import load_dataset ds = load_dataset("OpenMed/MedDialog") train = ds["train"] val = ds["validation"] print(train[0]["patient_message"]) print(train[0]["doctor_response"]) ``` ### With Prime Intellect RL Environment This dataset is used by the `maziyar/OpenMed_MedDialog` RL environment for training models via reinforcement learning with the following reward components: | Component | Weight | Description | |---|---|---| | Response Quality | 35% | Relevance, helpfulness, medical appropriateness | | Empathy & Communication | 25% | Patient-centered language, acknowledgment | | Medical Content | 20% | Addresses symptoms/concerns with relevant information | | Safety | 10% | Appropriate disclaimers, recommends professional consultation | | Fluency | 10% | Coherent, well-structured responses | ```bash prime env install maziyar/OpenMed_MedDialog ``` ## License This dataset is released under the [Apache License 2.0](https://www.apache.org/licenses/LICENSE-2.0), inherited from the source dataset [ruslanmv/ai-medical-chatbot](https://github.com/ruslanmv/ai-medical-chatbot). ## Limitations and Ethical Considerations - This dataset is intended for **research purposes only** and should not be used as a substitute for professional medical advice - Doctor responses in the source data vary in quality and may contain inaccuracies - The dataset reflects patterns from online medical Q&A platforms, which may not represent clinical best practices - Models trained on this data should include appropriate disclaimers about the limitations of AI-generated medical advice ## Citation If you use this dataset, please cite: ```bibtex @dataset{openmed_meddialog_2026, title={MedDialog: A Medical Dialogue Dataset for Clinical Response Generation}, author={OpenMed}, year={2026}, publisher={Hugging Face}, url={https://huggingface.co/datasets/OpenMed/MedDialog} } ``` ## Part of OpenMed This dataset is part of the [OpenMed](https://huggingface.co/OpenMed) collection of open medical NLP resources for research and development.

数据集显示名称:MedDialog 标签类型:中文语料库 许可协议:未知 媒体类型:医疗领域 论文链接:https://arxiv.org/pdf/2004.03329v2.pdf 发布日期:2020年 发布链接:https://github.com/UCSD-AI4H/Medical-Dialogue-System 发布机构:加利福尼亚大学圣地亚哥分校 标签:医疗、对话 --- # 数据集介绍 ## 简介 MedDialog(中文)数据集收录了医患双方的中文对话语料,共计110万组对话与400万条话语。该数据集仍在持续扩充,后续将新增更多对话内容。原始对话数据来源于好大夫在线(haodf.com),其全部版权归haodf.com所有。 ## 引文 @article{he2020meddialog, title={MedDialog: Two Large-scale Medical Dialogue Datasets}, author={He, Xuehai and Chen, Shu and Ju, Zeqian and Dong, Xiangyu and Fang, Hongchao and Wang, Sicheng and Yang, Yue and Zeng, Jiaqi and Zhang, Ruisi and Zhang, Ruoyu and others}, journal={arXiv preprint arXiv:2004.03329}, year={2020} } ## 数据集下载 :modelscope-code[]{type="git"}

提供机构:
maas
创建时间:
2026-01-25
搜集汇总
数据集介绍
MedDialog 数据集图片
背景与挑战
背景概述
MedDialog是一个大规模的中文医疗对话数据集,包含约110万条医生与患者之间的对话和400万条话语,数据来源于haodf.com。该数据集由加州大学圣地亚哥分校于2020年发布,主要用于医疗对话系统研究,具有丰富的真实世界医疗交流内容。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务