VD-Ref
收藏资源简介:
常规短语接地旨在将给定标题中提到的名词短语定位到其相应的图像区域,最近取得了巨大的成功。显然,唯一的名词短语基础不足以理解跨模式的视觉语言。在这里,我们还考虑代词来扩展任务。首先,我们构建一个短语数据集,该短语的基础包括名词短语和代词到图像区域。基于数据集,我们使用该行的最新文学模型来测试短语接地的性能。然后,我们使用核心参考信息来增强基线接地模型,这可能会帮助我们完成任务,并使用图卷积网络对核心参考结构进行建模。有趣的是,在我们的数据集上进行的实验表明,代词比名词短语更容易接地,其中可能的原因可能是这些代词的歧义要少得多。此外,我们具有核心参考信息的最终模型可以显著提高接地性能
Standard phrase grounding, which aims to locate noun phrases mentioned in a given caption to their corresponding image regions, has achieved remarkable recent success. Obviously, grounding only noun phrases is insufficient to comprehend cross-modal visual-language semantics. Herein, we extend this task by also incorporating pronouns. First, we construct a phrase grounding dataset where the grounding targets include both noun phrases and pronouns that need to be localized to image regions. Based on this dataset, we utilize state-of-the-art models from the field to evaluate the performance of phrase grounding. Subsequently, we enhance baseline grounding models with coreferential information to facilitate task accomplishment, and model coreferential structures using Graph Convolutional Networks (GCNs). Interestingly, experiments conducted on our dataset reveal that pronouns are easier to ground than noun phrases, with the underlying reason likely being that these pronouns have significantly less ambiguity. Furthermore, our final model equipped with coreferential information can substantially improve grounding performance.




