InfiMed-Foundation-1.7B, InfiMed-Foundation-4B
收藏资源简介:
InfiMed-Foundation数据集是两个医学专业的多模态大语言模型,旨在在医学应用中提供最先进的性能。该数据集结合了高质量的通用和医学多模态数据,并提出了一个新颖的五维质量评估框架来筛选高质量的多模态医学数据集。在持续预训练中,通过减少图像块数量和采用多模态序列打包来提高训练效率,从而能够整合大量的医学数据。此外,一个三阶段监督微调过程确保了复杂医学任务的有效知识提取。该数据集旨在解决医学领域中通用多模态大语言模型缺乏专业知识的问题,并通过高质量的数据集、高效的训练策略和领域特定知识的有效提取,为医疗保健领域提供了更可靠和有效的AI驱动解决方案。
The InfiMed-Foundation dataset underpins a multimodal large language model (LLM) designed for two medical specialties, targeting state-of-the-art performance in medical applications. It integrates high-quality general and medical multimodal data, and introduces a novel five-dimensional quality assessment framework to filter high-quality multimodal medical datasets. For continuous pre-training, training efficiency is improved by reducing the quantity of image patches and adopting multimodal sequence packing, thereby enabling the integration of massive volumes of medical data. Additionally, a three-stage supervised fine-tuning pipeline ensures effective knowledge extraction for complex medical tasks. This dataset aims to address the shortage of specialized domain knowledge in general-purpose multimodal large language models for the medical field, and delivers more reliable and effective AI-driven solutions for the healthcare sector via high-quality multimodal data, efficient training strategies, and effective extraction of domain-specific knowledge.




