BiMed-V
收藏资源简介:
BiMed-V是一个综合的双语(阿拉伯语-英语)多模态医疗数据集,包含160万条样本,旨在提升医疗图像与文本的对齐和多模态理解能力。数据集内容丰富,涵盖多种公开数据集和自定义数据,支持多种医疗图像模态,如胸部X光、CT、MRI等。数据集的创建过程包括翻译和专家验证,确保了数据的质量和临床相关性。该数据集主要应用于医疗领域的多模态任务,如报告生成、图像问答等,旨在解决多语言医疗AI模型的需求问题。
BiMed-V is a comprehensive bilingual (Arabic-English) multimodal medical dataset consisting of 1.6 million samples, designed to enhance the alignment between medical images and text as well as multimodal understanding capabilities. The dataset encompasses diverse content including multiple public datasets and custom-curated data, and supports various medical image modalities such as chest X-rays, CT scans, and MRI scans. The development process of the dataset involves translation and expert validation, which ensures data quality and clinical relevance. This dataset is primarily utilized for multimodal medical tasks including report generation and visual question answering (VQA), with the goal of addressing the demand for multilingual medical AI models.




