FIOVA
收藏资源简介:
FIOVA数据集由南洋理工大学和中国科学院自动化研究所等机构共同创建,包含3002个长视频序列,平均时长为33.6秒,涵盖38个不同的主题。每个视频由五名不同的标注者进行详细标注,生成4到15倍于现有基准的描述长度,旨在建立一个全面的人类理解基线。数据集的创建过程严格遵循标准化指南,确保描述的准确性和一致性。FIOVA数据集主要用于评估和比较大型视觉语言模型与人类在视频理解任务中的表现,旨在解决视频描述任务中的复杂时空关系理解问题。
The FIOVA dataset was co-created by institutions including Nanyang Technological University and the Institute of Automation of the Chinese Academy of Sciences. It comprises 3002 long video sequences with an average duration of 33.6 seconds, covering 38 distinct topics. Each video was meticulously annotated by five separate annotators, generating descriptive content that is 4 to 15 times longer than that of existing benchmark datasets, with the goal of establishing a comprehensive human understanding baseline. The dataset's development process strictly follows standardized guidelines to ensure the accuracy and consistency of the annotations. The FIOVA dataset is primarily used to evaluate and compare the performance of large vision-language models and humans in video understanding tasks, aiming to address the challenge of understanding complex spatio-temporal relationships in video captioning tasks.

- 1Can LVLMs Describe Videos like Humans? A Five-in-One Video Annotations Benchmark for Better Human-Machine Comparison南洋理工大学、中国科学院自动化研究所、中国科学院大学、北京理工大学、东南大学、北京科技大学 · 2024年



