CODIS
收藏资源简介:
CODIS数据集由清华大学计算机科学与技术系创建,旨在评估多模态大型语言模型在依赖上下文的视觉理解能力。该数据集包含377张图片,每张图片都具有内在的模糊性,需要额外的自由格式文本上下文才能准确解读。数据集设计强调了上下文在视觉理解中的重要性,特别是在解决图像模糊性方面。CODIS数据集的应用领域包括视觉问答和视觉推理,旨在通过上下文依赖的方式提高模型的视觉理解能力,解决现有模型在处理视觉信息时对上下文理解的不足。
The CODIS dataset was developed by the Department of Computer Science and Technology, Tsinghua University, with the goal of evaluating the context-dependent visual understanding abilities of multimodal large language models. This dataset comprises 377 images, each of which possesses inherent ambiguity and demands additional free-form textual context for accurate interpretation. The design of the CODIS dataset highlights the critical role of context in visual understanding, particularly in resolving image ambiguity. Application scenarios of the CODIS dataset cover visual question answering and visual reasoning; it aims to improve the visual understanding capabilities of models through context-dependent methods, addressing the insufficiency of existing models in contextual comprehension when processing visual information.

- 1CODIS: Benchmarking Context-Dependent Visual Comprehension for Multimodal Large Language Models清华大学计算机科学与技术系 · 2024年



