williamliu/ChiMed-VL
收藏资源简介:
# ChiMed-VL Dataset ## ChiMed-VL-Alignment dataset ## ChiMed-VL-Alignment consists of 580,014 image-text couplings, each pair falling into one of two categories: context information of an image or descriptions of an image. The context category contains 167M tokens, presenting a median text length of 435 (Q1: 211, Q3: 757). Conversely, descriptions, more concise and image-specific, contain inline descriptions and captions. They comprise 63M tokens, with median lengths settling at 59 (Q1: 45, Q3: 83). ## ChiMed-VL-Instruction dataset ## ChiMed-VL-Instruction comprises 469,441 question-answer pairs. Within this subset, the questions section contains 10M tokens with a median length of 20 (Q1: 16, Q3: 25), posing a concise inquiry reflective of medical queries. The answers consist of 13M tokens with a median length slightly longer at 22 (Q1: 12, Q3: 34), providing clear, direct, and informative responses.
# ChiMed-VL 数据集 ## ChiMed-VL-Alignment 数据集 ## ChiMed-VL-Alignment 数据集包含580,014组图文配对样本,每一组均属于以下两类之一:图像上下文信息,或图像描述文本。其中上下文信息类别的文本总计包含1.67亿个Token(Token),文本长度的中位数为435(四分位距Q1:211,Q3:757)。与之相对的图像描述文本类别更为简洁且贴合图像主题,涵盖内嵌描述与图像字幕,总计包含6300万个Token,文本长度中位数为59(Q1:45,Q3:83)。 ## ChiMed-VL-Instruction 数据集 ## ChiMed-VL-Instruction 数据集包含469,441组问答配对样本。该子集中的问题文本总计包含1000万个Token,文本长度中位数为20(Q1:16,Q3:25),问题表述简洁,贴合医疗查询场景。回答文本总计包含1300万个Token,文本长度中位数略长,为22(Q1:12,Q3:34),所生成的回答清晰直白且信息详实。
ChiMed-VL Dataset 概述
ChiMed-VL-Alignment 数据集
- 数据量: 包含580,014个图像-文本对。
- 内容分类: 分为两类,即图像的上下文信息和图像描述。
- 上下文信息:
- 文本量: 包含167M tokens。
- 文本长度: 中位数为435,第一四分位数(Q1)为211,第三四分位数(Q3)为757。
- 图像描述:
- 文本量: 包含63M tokens。
- 文本长度: 中位数为59,第一四分位数(Q1)为45,第三四分位数(Q3)为83。
ChiMed-VL-Instruction 数据集
- 数据量: 包含469,441个问题-答案对。
- 问题部分:
- 文本量: 包含10M tokens。
- 文本长度: 中位数为20,第一四分位数(Q1)为16,第三四分位数(Q3)为25。
- 答案部分:
- 文本量: 包含13M tokens。
- 文本长度: 中位数为22,第一四分位数(Q1)为12,第三四分位数(Q3)为34。




