NJU-LINK/ViDiC-1K
收藏资源简介:
ViDiC-1K是一个视频差异描述(Video Difference Captioning)的基准数据集,包含1000个精心挑选的视频对,标注了超过4000个比较检查项。该数据集旨在评估多模态大语言模型(MLLMs)在视频对之间提供细粒度相似性和差异性描述的能力。数据集的特点包括首次视频差异描述基准、双检查表评估、可扩展的LLM-as-a-Judge评估协议等。视频时长主要在2-12秒之间,数据来源包括公共数据集和自生成的合成数据。
ViDiC-1K is a benchmark dataset for Video Difference Captioning, comprising 1,000 curated video pairs annotated with over 4,000 comparative checklist items. The dataset is designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to provide fine-grained descriptions of similarities and differences between video pairs. Key features include the first video difference captioning benchmark, dual-checklist evaluation, and a scalable LLM-as-a-Judge evaluation protocol. Video durations are primarily 2-12 seconds, with data sourced from public datasets and self-generated synthetic data.




