Cops-Ref
收藏资源简介:
Cops-Ref数据集由香港大学创建,专注于组合指称表达理解,旨在通过复杂的语言表达识别图像中的特定对象。该数据集包含148,712条表达式,基于真实世界图像,强调视觉真实性和语义丰富性。数据集的创建过程中,设计了六种逻辑形式,灵活结合丰富的视觉信息生成具有不同组合性的表达式。Cops-Ref的应用领域包括视觉问答和视觉对话,旨在解决模型在复杂视觉场景中理解和定位对象的问题。
The Cops-Ref dataset, created by The University of Hong Kong, focuses on compositional referring expression comprehension, with the goal of identifying specific objects in images via complex linguistic expressions. This dataset comprises 148,712 expressions, which are developed based on real-world images and prioritize visual authenticity and semantic richness. During the dataset construction process, six logical forms were designed to flexibly combine rich visual information and generate expressions with diverse levels of compositionality. The application domains of Cops-Ref include visual question answering (VQA) and visual dialogue, aiming to address the challenges of models understanding and localizing objects in complex visual scenes.




