TCM-Ladder
收藏资源简介:
TCM-Ladder是第一个专门为评估大型TCM语言模型设计的多模态QA数据集,涵盖中医多个核心学科,包括基础理论、诊断学、方剂学、内科学、外科学、生药学和儿科学。除了文本内容,TCM-Ladder还包含图像和视频等多种模态。数据集通过自动和人工过滤结合的方式构建,总计包含52,000多个问题,包括单选题、多选题、填空题、诊断对话和视觉理解任务。
TCM-Ladder is the first multimodal question answering (QA) dataset specifically designed for evaluating large Traditional Chinese Medicine (TCM) language models. It covers multiple core TCM disciplines including basic theory, diagnostics, formulology, internal medicine, surgery, pharmacognosy, and pediatrics. In addition to textual content, TCM-Ladder also supports multiple modalities such as images and videos. The dataset is constructed through a combination of automatic and manual filtering, and contains a total of over 52,000 questions including single-choice questions, multiple-choice questions, fill-in-the-blank questions, diagnostic dialogues, and visual understanding tasks.
TCM-Ladder 数据集概述
数据集简介
- 名称:TCM-Ladder
- 领域:传统中医(TCM)
- 类型:多模态问答数据集
- 目的:评估大型中医语言模型在真实任务中的表现
核心特点
- 多模态性:包含文本、图像、视频等多种数据形式
- 广泛覆盖:涵盖中医核心学科领域
- 基础理论
- 诊断学
- 方剂学
- 内科学
- 外科学
- 生药学
- 儿科学
- 任务多样性:包含6种任务类型
- 单选题(基础知识识别)
- 多选题(复杂概念整合推理)
- 长形式诊断问答(临床推理)
- 填空题(生成准确性和上下文理解)
- 基于图像的理解任务(多模态推理)
- 音频/视频资源(支持多模态模型开发)
数据规模
- 总问题量:52,000+
- 构建方法:自动化与人工筛选结合
评估方法
- Ladder-Score:专门设计的中医问答评估方法
- 评估术语使用
- 评估语义表达质量
实验验证
- 对比模型:
- 9个最先进的通用领域LLM
- 5个领先的中医专用LLM
- 评估维度:
- 单/多选题表现
- 中药材相关问题表现
- 舌像图像问题表现
可用资源
- 数据集访问:https://tcmladder.com 或 https://54.211.107.106
- 持续更新:是




