YFCC15M
收藏资源简介:
YFCC15M是由北京航空航天大学和SenseTime Research合作创建的大型图像-文本数据集,包含15388848对图像和文本描述。该数据集通过精细的过滤策略,提高了数据质量,主要用于评估和分析对比语言-图像预训练(CLIP)模型的性能。YFCC15M支持多种视觉任务,如零样本识别和图像分类,旨在通过高质量的数据提升模型的泛化能力和训练效率。
YFCC15M is a large-scale image-text dataset co-created by Beihang University and SenseTime Research, containing 15,388,848 pairs of images and their corresponding textual descriptions. This dataset adopts a rigorous filtering strategy to improve data quality, and is primarily used for evaluating and analyzing the performance of Contrastive Language-Image Pre-training (CLIP) models. YFCC15M supports multiple visual tasks such as zero-shot recognition and image classification, aiming to enhance the generalization ability and training efficiency of models via high-quality data.

- 1Democratizing Contrastive Language-Image Pre-training: A CLIP Benchmark of Data, Model, and Supervision北京航空航天大学 · 2022年



