OpenNLPLab/FAVDBench
收藏资源简介:
在CVPR2023中我们提出了精细化音视频描述任务(Fine-grained Audible Video Description, FAVD)该任务旨在提供有关可听视频的详细文本描述,包括每个对象的外观和空间位置、移动对象的动作以及视频中的声音。我们同是也为社区贡献了第一个精细化音视频描述数据集FAVDBench。对于每个视频片段,我们不仅提供一句话的视频概要,还提供4-6句描述视频的视觉细节和1-2个音频相关描述,且所有的标注都有中英文双语。
At CVPR 2023, we introduced the Fine-grained Audible Video Description (FAVD) task, which aims to generate detailed textual descriptions for audible videos. The task covers the appearance and spatial positions of each object, the actions of moving objects, and the sounds present in the video. Additionally, we contributed the first fine-grained audible video description dataset, FAVDBench, to the research community. For each video clip, our annotations include not only a one-sentence video summary, but also 4 to 6 descriptions of the video's visual details and 1 to 2 audio-related descriptions. All annotations are provided in both Chinese and English.




