TCM-Ladder
收藏资源简介:
TCM-Ladder是一个大规模的多模态数据集,旨在为中医药领域的大型语言模型提供训练和评估。该数据集涵盖了中医药的多个子学科,包括基本理论、诊断、药方、药理学等,并融合了文本、图像、音频和视频等多种数据类型。数据集的建设过程中,收集了超过52,000个问题,包括单选题、多选题、填空题、诊断对话和视觉理解任务等。所有文本和视觉数据均由认证的中医药从业者独立审查和验证,以确保准确性和临床相关性。
TCM-Ladder is a large-scale multimodal dataset developed to support the training and evaluation of large language models within the traditional Chinese medicine (TCM) domain. This dataset spans multiple sub-disciplines of TCM, including basic theories, clinical diagnosis, prescriptions, pharmacology and other related fields, and integrates diverse data modalities such as text, images, audio and video. Over 52,000 questions were collected during the construction of this dataset, covering single-choice questions, multiple-choice questions, fill-in-the-blank questions, diagnostic dialogues and visual understanding tasks. All text and visual data were independently reviewed and validated by certified TCM practitioners to ensure their accuracy and clinical relevance.



