FIBER
收藏资源简介:
FIBER是一个细粒度的视频-文本检索基准数据集,由南京大学和上海人工智能实验室联合创建。该数据集包含1000个视频,来源于FineAction数据集,每个视频都配备了高质量的人工注释,涵盖了视频的静态场景、动态动作、拍摄风格等多方面细节。数据集的注释分为空间和时间两部分,能够独立评估模型的空间和时间偏差。FIBER旨在解决现有视频检索基准在细粒度检索能力评估上的不足,特别适用于评估多模态大语言模型在视频检索任务中的表现。
FIBER is a fine-grained video-text retrieval benchmark dataset jointly created by Nanjing University and Shanghai AI Laboratory. This dataset includes 1000 videos sourced from the FineAction dataset, each paired with high-quality manual annotations covering multiple details such as the video's static scenes, dynamic actions, shooting styles and other aspects. The annotations of the dataset are divided into spatial and temporal parts, which can independently evaluate the spatial and temporal biases of models. FIBER aims to address the shortcomings of existing video retrieval benchmarks in evaluating fine-grained retrieval capabilities, and is particularly suitable for assessing the performance of multimodal large language models in video retrieval tasks.




