DataComp for VLMs (DCVLM)
收藏资源简介:
DCVLM是由多个研究机构联合构建的综合性视觉-语言模型训练基准,旨在系统化研究数据策展策略。该数据集整合了160个公开来源,涵盖图文对、多模态交错文档、纯文本及指令微调数据四种类型,总计包含6万亿多模态tokens,数据来源高度异构且覆盖超过20种语言。其构建过程通过标准化数据池汇集与去污染处理实现,确保训练与评估无重叠。该数据集主要应用于视觉-语言模型的预训练优化研究,重点解决多源数据混合策略、规模扩展效应及跨领域评估标准化等核心问题。
DCVLM is a comprehensive visual-language model training benchmark jointly constructed by multiple research institutions, aiming to systematically investigate data curation strategies. This dataset incorporates 160 public sources, encompassing four data types: image-text pairs, multimodal interleaved documents, plain text, and instruction-tuning data, with a total of 6 trillion multimodal tokens. Its data sources are highly heterogeneous and span over 20 languages. The construction of DCVLM adopts standardized dataset pool aggregation and decontamination processing to eliminate any overlap between training and evaluation datasets. This dataset is primarily utilized for research on pretraining optimization of visual-language models, focusing on addressing core challenges including multi-source data mixing strategies, scaling effects, and cross-domain evaluation standardization.




