Alibaba-DAMO-Academy/InterVBench
收藏资源简介:
LV-Bench是一个精心策划的基准测试数据集,包含1000个时长一分钟的视频,旨在评估长时程生成。视频来源于DanceTrack、GOT-10k、HD-VILA-100M和ShareGPT4V,类别分布大约为67%以人类为中心、17%以动物为中心和16%以环境为中心的镜头。每个源视频被分割成2-3秒的片段,并使用GPT-4o生成字幕,随后在每个阶段(来源选择、分割、字幕审核)都经过人工验证以确保质量。该基准测试分为80/20的训练-评估分割,并结合VDE套件和标准VBench分数,为时间一致性提供全面的压力测试。
LV-Bench is a curated benchmark of 1,000 minute-long videos targeted at evaluating long-horizon generation. Videos are sourced from DanceTrack, GOT-10k, HD-VILA-100M, and ShareGPT4V, yielding a class distribution of roughly 67% human-focused, 17% animal-focused, and 16% environment-focused footage. Each source video is broken into 2–3 second segments and captioned with GPT-4o, followed by human validation at every stage (sourcing, chunking, caption review) to maintain quality. The benchmark is divided into an 80/20 train-eval split and pairs the VDE suite with standard VBench scores, providing a comprehensive stress test for temporal coherence.




