BennoKrojer/ImageCoDe
收藏资源简介:
ImageCoDe是一个视觉与语言理解的基准测试,要求理解语用学、时间性、长描述和视觉细微差别。该任务要求根据详细描述从10个最小对比图像中检索目标图像。数据集包含21K描述和94K图像,主要基于视频数据集的帧。
ImageCoDe is a benchmark for vision-and-language understanding that requires comprehension of pragmatics, temporality, long-form descriptions and visual nuances. The task mandates retrieving the target image from 10 minimally contrasting images based on a detailed description. The dataset comprises 21,000 descriptions and 94,000 images, which are primarily derived from frames of video datasets.
ImageCoDe数据集概述
数据集描述
- 任务类型:视觉与语言基准,要求在给定详细描述的情况下,从10张最小对比度的图像中检索目标图像。
- 数据内容:包含21,000个描述和94,000张图像,图像主要基于视频数据集的帧。
数据集结构
数据实例
每个实例包含以下信息:
- 描述
- 对应的图像集名称
- 目标图像索引
示例:
{"image_set": "video-storytelling-videowedding_de8dLXvgV-I-shot6_0", "image_index": "8", "description": "The flowers the woman in the teal strapless dress is carrying are completely obscured by the man in the black shirts head. "}
数据分割
| 数据集分割 | 描述数量 |
|---|---|
| 训练集 | 16,594 |
| 验证集 | 2,302 |
| 测试集 | 2,306 |
数据集创建
精选理由
ImageCoDe旨在揭示近期视觉与语言模型在处理复杂语言和精细视觉表示方面的弱点。此外,该数据集提供了大量实用的示例,适合研究语用学。




