遇见数据集

HL Dataset: Visually-grounded Description of Scenes, Actions and Rationales

收藏
Zenodo2024-02-28 更新2026-05-26 收录
官方服务:

资源简介:

Current captioning datasets focus on object-centric captions, describing the visible objects in the image, often ending up stating the obvious (for humans), e.g. "people eating food in a park". Although these datasets are useful to evaluate the ability of Vision & Language models to recognize and describe visual content, they do not support controlled experiments involving model testing or fine-tuning, with more high-level captions, which humans find easy and natural to produce. For example, people often describe images based on the type of scene they depict ("people at a holiday resort") and the actions they perform ("people having a picnic"). Such concepts are based on personal experience and contribute to forming common sense assumptions. We present the High-Level Dataset, a dataset extending 14997 images from the COCO dataset, aligned with a new set of 134,973 human-annotated (high-level) captions collected along three axes: scenes, actions and rationales. We further extend this dataset with confidence scores collected from an independent set of readers, as well as a set of narrative captions generated synthetically, by combining each of the three axes. We describe this dataset and analyse it extensively. We also present baseline results for the High-Level Captioning task.

当前的图像描述数据集多以物体为中心,仅对图像中可见的实体进行描述,最终往往产出人类视角下显而易见的内容,例如“人们在公园进食”。尽管此类数据集可用于评估视觉语言(Vision & Language)模型识别与描述视觉内容的能力,但它们无法支撑围绕人类易于自然生成的高级图像描述开展的可控实验,难以用于模型测试或微调。例如,人类通常会基于图像所展现的场景类型(如“度假胜地中的人群”)以及人物行为(如“人们正在野餐”)来描述图像。此类概念基于个人经验,有助于构建常识性认知假设。本文提出高级图像描述数据集(High-Level Dataset),该数据集以COCO数据集的14997张图像为基础进行拓展,为其匹配了134973条人工标注的高级图像描述,标注维度涵盖场景、行为与动机三大类。本研究进一步为该数据集补充了由独立标注读者标注的置信度分数,同时通过组合上述三大标注维度,合成生成了一批叙事性图像描述,一并纳入数据集。本文对该数据集进行了详细介绍与全面分析。此外,本文还给出了高级图像描述任务的基线实验结果。

创建时间:
2024-02-28
二维码
社区交流群
二维码
科研交流群
商业服务