MS-COCO, Flickr30k, MS-COCO-FG, Flickr30k-FG
收藏资源简介:
本文涉及的数据集包括MS-COCO、Flickr30k及其增强版本MS-COCO-FG和Flickr30k-FG,这些数据集主要用于图像-文本检索任务。MS-COCO数据集包含123,287张图片和616,435个描述,而Flickr30k包含31,783张图片和158,915个描述。增强版本的数据集通过添加额外的上下文细节来提高描述的详细程度。这些数据集的创建旨在通过提供更精细的图像描述来改善图像-文本检索模型的性能,特别是在概念粒度方面。数据集的应用领域主要是信息检索和图像-文本匹配,旨在解决现有基准数据集在细节描述和评估方法上的不足。
The datasets covered in this paper include MS-COCO, Flickr30k, and their enhanced variants MS-COCO-FG and Flickr30k-FG, which are primarily utilized for image-text retrieval tasks. MS-COCO consists of 123,287 images paired with 616,435 captions, while Flickr30k contains 31,783 images and 158,915 captions. The enhanced versions of these datasets improve the detail level of image descriptions by adding additional contextual details. These datasets are developed to enhance the performance of image-text retrieval models, especially in terms of conceptual granularity, by providing more fine-grained image descriptions. The main application fields of these datasets are information retrieval and image-text matching, aiming to address the shortcomings of existing benchmark datasets in detail description and evaluation methods.

- 1Assessing Brittleness of Image-Text Retrieval Benchmarks from Vision-Language Models Perspective阿姆斯特丹大学 · 2024年



