HarMoE Harmonized Dataset
收藏资源简介:
HarMoE Harmonized Dataset是由伯尔尼大学等机构整合多源胸部X光影像构建的统一预训练数据集,旨在突破单源依赖,提升模型在异构数据下的泛化能力。该数据集融合了MIMIC-CXR、CheXpert、ChestX-ray14和PadChest四个经典数据集,总计约87.3万张图像,覆盖229种疾病标签。构建过程通过大语言模型统一报告标签、构建全局疾病词汇表并采用三态编码,有效避免了未标注类别的假阴性问题。该数据集主要用于胸部X光视觉-语言预训练,解决多源数据中标签体系不一致与领域偏移带来的表示纠缠问题,显著提升零样本分类与跨域迁移性能。
The HarMoE Harmonized Dataset is a unified pre-training dataset developed by the University of Bern and other institutions by integrating multi-source chest X-ray images. It aims to break through the reliance on single-source data and enhance the generalization capability of models under heterogeneous data conditions. This dataset integrates four classic datasets including MIMIC-CXR, CheXpert, ChestX-ray14 and PadChest, with a total of approximately 873,000 images covering 229 disease labels. During its construction, large language models (LLMs) are used to unify report labels, build a global disease vocabulary and adopt three-state encoding, which effectively avoids false negative problems caused by unannotated categories. This dataset is mainly applied to chest X-ray vision-language pre-training, addressing the issue of representation entanglement arising from inconsistent label systems and domain shift in multi-source data, and significantly improving the performance of zero-shot classification and cross-domain transfer.




