futo-org/coco-captions-qwen3vl
收藏资源简介:
这是一个高质量、面向检索的文档式描述数据集,基于COCO(Karpathy训练分割)的113,287张图像。它提供丰富、接地气、属性密集的描述,旨在作为搜索文档(60-150个单词),而非单行描述。数据集覆盖COCO 2014训练+验证池中的图像(标准分割用于描述/检索工作),未包含保留的Karpathy验证/测试集(各5,000张)。每张图像通过Qwen3-VL-30B-A3B-Instruct模型生成16个候选描述,然后使用Qwen3-VL-Reranker-8B模型对所有候选进行排名,保留前2个最佳描述。原始COCO人类描述(每张图像5个)作为生成提示提供,但未包含在数据集中。所有113,287张图像均被覆盖,确保数据完整性。
High-quality, retrieval-oriented document-style captions for the 113k COCO (Karpathy train split) images: rich, grounded, attribute-dense descriptions intended as search documents (60–150 words), not one-line captions. This dataset covers the Karpathy training split: 113,287 images drawn from the COCO 2014 train+val pool (the standard split for caption/retrieval work; the held-out Karpathy val/test 5k+5k are deliberately not included). For each image, 16 candidate captions were generated with Qwen3-VL-30B-A3B-Instruct, then ranked with Qwen3-VL-Reranker-8B and the top-2 were kept. Original COCO human captions were used as hints but are not included. All 113,287 images are covered.




