MTCMB datasets
收藏资源简介:
MTCMB (Multi-Task benchmark dataset for evaluating large language models in Traditional Chinese Medicine) is a comprehensive evaluation suite designed to assess LLMs on TCM knowledge, reasoning, and safety. Co-developed with certified TCM practitioners, the dataset is curated from standardized examinations, authoritative textbooks, official pharmacopoeia, and academic competition datasets. The benchmark comprises 12 sub-datasets spanning five task categories: Knowledge QA: Assesses factual recall. TCM-ED-A (1,200 samples): Multi-choice QA for 12 clinical disciplines. Sourced from the National TCM Attending Physician Examination Question Bank. TCM-ED-B (4,800 samples): Full-length multi-choice exams. Sourced from the National TCM Practitioner Examination Question Bank. TCM-FT (100 samples): Open-ended Q&A. Sourced from the "TCM Question and Answer Bank". Language Understanding: Evaluates entity recognition, dialogue reconstruction, and literature comprehension. TCMeEE (100 samples): Entity extraction (symptoms, signs, formulas). 95% sourced from the TCM Think Tank (zhongyigen.com) and 5% from medical cases submitted by professional practitioners. TCM-CHGD (100 samples): Structured medical record generation from doctor-patient dialogues. Sourced from the National TCM Practitioner (Assistant) Qualification Examination Practical Skills Question Bank. TCM-LitData (100 samples): QA based on classical TCM text excerpts. Sourced from the Aliyun Tianchi TCM Literature Question Generation Dataset. Diagnostic Reasoning: Evaluates clinical inference capabilities. TCM-MSDD (100 samples): Multi-label disease/syndrome classification based on natural language symptom descriptions. Sourced from the CCL25-Eval Task 9 Subtask 1 Dataset. TCM-Diagnosis (200 samples): Generation of complete diagnoses (TCM disease and syndrome). Sourced from the 13th Five-Year Plan TCM Textbooks (Internal Medicine, Surgery, Gynecology, and Pediatrics). Prescription Recommendation: Validates treatment logic. TCM-PR (100 samples): Formula prediction from case descriptions. Sourced from the CCL25-Eval Task 9 Subtask 2 Dataset. TCM-FRD (200 samples): Generation of detailed treatment plans (methods, formula name, and herbs). Sourced from the 13th Five-Year Plan TCM Textbooks (Internal Medicine, Surgery, Gynecology, and Pediatrics). Safety Evaluation: Tests knowledge of medical contraindications. TCM-SE-A (50 samples): Fill-in-the-blank questions regarding contraindications. Sourced from the Pharmacopoeia of the People's Republic of China (2020 Edition). TCM-SE-B (50 samples): Multiple-choice questions on medication and safety. Sourced from the Pharmacopoeia of the People's Republic of China (2020 Edition).



