arkarchanmyae/medical-question-answering-datasets
收藏资源简介:
这是一个用于问答任务的医疗领域数据集,包含多个子集,涵盖临床、医疗保健等主题。数据集由多个配置组成,包括all-processed、chatdoctor_healthcaremagic、chatdoctor_icliniq以及medical_meadow系列(如cord19、health_advice、medical_flashcards、mediqa、medqa、mmmlu、pubmed_causal、wikidoc和wikidoc_patient_information)。每个配置均包含instruction、input和output三个字符串特征,用于训练问答模型。数据集总计包含超过246,678个训练示例,数据量约为276,980,695字节,语言为英语,许可证为MIT。
This is a medical domain dataset for question-answering tasks, comprising multiple subsets covering topics such as clinical and healthcare. The dataset consists of several configurations, including all-processed, chatdoctor_healthcaremagic, chatdoctor_icliniq, and the medical_meadow series (e.g., cord19, health_advice, medical_flashcards, mediqa, medqa, mmmlu, pubmed_causal, wikidoc, and wikidoc_patient_information). Each configuration features three string attributes: instruction, input, and output, designed for training question-answering models. The dataset includes over 246,678 training examples in total, with a data size of approximately 276,980,695 bytes. The language is English, and it is licensed under MIT.



