Localized Narratives
收藏资源简介:
该数据集包含了84.9万张图片,每张图片都附有与描述中的每个单词对齐的鼠标轨迹注释,展现了视觉注意力与语言描述之间的关系。此外,该项目涉及了156位专业标注人员,在名词和动词上的语义准确度达到了98.0%。同时,还分析了鼠标轨迹与物体位置之间的准确性。该数据集的规模达到了84.9万张已标注的图片,其任务是对视觉语言模型中的视觉注意力进行鼠标轨迹注释。
This dataset contains 849,000 images, each paired with mouse trajectory annotations aligned to every word in its accompanying description, which elucidates the relationship between visual attention and linguistic descriptions. The project recruited 156 professional annotators, achieving a semantic accuracy of 98.0% for both nouns and verbs. Additionally, the alignment accuracy between mouse trajectories and object positions was evaluated. With a total of 849,000 annotated images, this dataset is dedicated to the task of annotating visual attention in vision-language models via mouse trajectories.




