Monor/hwtcm
收藏资源简介:
该数据集可用于评估大型语言模型(LLM)在传统中医领域的能力,包含多选题、单选题和判断题。数据集提供了每种题型的示例,并展示了不同模型在这些题型上的准确率。此外,还提到了一个正在训练中的中医领域的大型语言模型Canggong-14b-chat。
This dataset can be used to evaluate the traditional Chinese medicine capabilities of LLM and contains multiple-choice, multiple-answers and true/false questions. The dataset provides examples for each type of question and shows the accuracy of different models on these question types. Additionally, it mentions a large language model in the field of traditional Chinese medicine, Canggong-14b-chat, which is still in training.
数据集概述
基本信息
- 许可证: MIT
- 任务类别: 问答
- 语言: 中文
- 标签: 医疗, 中医, 传统中医, 评估, 基准测试, 测试
数据集描述
该数据集用于评估大型语言模型(LLM)在传统中医方面的能力,包含多选题、多答案题和判断题。
数据示例
多答案题
json [ { "instruction": "请阅读以下中医评级考试多项选择题,并选出最合适的答案。", "input": "便秘的预防调护应注意 A.保持心情舒畅 B.少吃辛辣刺激性食物 C.适当摄入油脂 D.积极治疗肛门直肠疾病 E.按时登厕", "output": "ABCDE" } ]
单选题
json [ { "instruction": "以下是关于中医考试的选择题,请认真作答并选出正确答案。", "input": "患者,男,50岁。眩晕欲仆,头摇而痛,项强肢颤,腰膝疫软,舌红苔薄白,脉弦有力。其病机是 A.肝阳上亢 B.肝肾阴虚 C.肝阳化风 D.阴虚风动 E.肝血不足", "output": "C" } ]
判断题
json [ { "instruction": "请仔细阅读以下中医学测验判断题,随后进行正确判断。", "input": "秦医医和提出了“六气病源说”。", "output": "正确" }, { "instruction": "下面是中医学考试的判断题,请认真阅读并作出正确判断。", "input": "中风中经络邪盛时也可出现神志改变", "output": "错误" } ]
模型准确率基准测试
| 模型名称 | 单选题准确率 | 多答案题准确率 | 判断题准确率 |
|---|---|---|---|
| llama3:8b | 21.94% | 17.71% | 46.56% |
| phi3:14b-instruct | 26.93% | 1.04% | 38.93% |
| aya:8b | 17.85% | 1.04% | 34.35% |
| mistral:7b-instruct | 21.76% | 2.08% | 48.09% |
| qwen1.5-7b-chat | 51.35% | 13.54% | 46.56% |
| qwen1.5-14b-chat | 69.94% | 78.12% | 31.30% |
| huangdi-13b-chat | 21.73% | 45.83% | 0.00% |
| canggong-14b-chat(SFT) | 55.98% | 4.17% | 23.66% |
| canggong-14b-chat(DPO) | 72.33% | 2.08% | 45.80% |




