Cola
收藏资源简介:
Cola数据集是由波士顿大学的研究团队开发的,用于评估大型视觉语言模型在组合对象定位与属性方面的能力。该数据集包含约1236个组合查询,涉及168个对象和197个属性,分布在约30K张图像上。创建过程中,模型必须从包含相同对象和属性但配置错误的干扰图像中检索出正确配置的图像。Cola数据集主要应用于测试和改进视觉语言模型在组合推理方面的性能,特别是在辅助代理等实际应用中,需要理解细粒度差异的场景。
The Cola dataset was developed by a research team at Boston University to assess the compositional object localization and attribute comprehension capabilities of large vision-language models. This dataset includes approximately 1,236 compositional queries, covering 168 object categories and 197 attribute categories, spanning roughly 30,000 images. The dataset’s task framework requires models to retrieve correctly configured images from distractor images that share identical objects and attributes but feature misconfigured setups. The Cola dataset is primarily utilized to test and enhance the compositional reasoning performance of vision-language models, particularly in practical scenarios such as AI agent applications that demand fine-grained understanding of subtle differences.

- 1COLA: A Benchmark for Compositional Text-to-image Retrieval波士顿大学 · 2023年



