Movie Facts and Fibs (MF2)
收藏资源简介:
MF2数据集是一个用于评估模型对完整电影(时长50-170分钟)理解程度的新基准。该数据集包含超过50部完整长度的、开放许可的电影,每部电影都配有一套手动构建的声明对——一个真实的(事实)和一个看似合理但错误的(谎言),共计超过850对。这些声明针对电影中的核心叙事元素,如角色动机和情绪、因果链和事件顺序,并引用人们无需重看电影就能回忆起的重要时刻。与多项选择题格式不同,我们采用二元声明评估协议:对于每对声明,模型必须正确识别出真实和错误的声明。这减少了答案排序等偏差,并能够更精确地评估推理能力。我们的实验表明,无论是开放权重还是封闭的顶级模型,其性能都远低于人类,突显了人类在记忆和推理关键叙事信息方面的优越能力,这是当前视觉-语言模型所缺乏的。
The MF2 dataset is a novel benchmark for evaluating models' understanding of full-length feature films with durations ranging from 50 to 170 minutes. This dataset includes over 50 full-length, openly licensed films, each paired with a manually constructed statement pair: one factual (true) statement and one plausible but incorrect (false/lie) statement, totaling more than 850 pairs. These statements target core narrative elements in the films, such as character motivations and emotions, causal chains and event sequences, and reference key moments that viewers can recall without re-watching the films. In contrast to multiple-choice question formats, we adopt a binary statement evaluation protocol: for each pair of statements, the model must correctly identify which one is true and which one is false. This reduces biases such as answer ordering, and enables more precise evaluation of model reasoning capabilities. Our experiments show that both open-weight and closed-weight state-of-the-art models perform far worse than humans, highlighting the superior human ability to memorize and reason about key narrative information, which current vision-language models lack.
数据集概述
基本信息
- 名称: sardinelab/MF2
- 许可证: CC BY-NC-SA 4.0
- 任务类别: 视觉问答 (Visual Question Answering)
- 语言: 英语 (en)
数据集特点
- 标签:
- 长电影理解 (Long Movie Understanding)
- 多模态 (Multimodal)
- 规模类别: 小于1K样本 (n<1K)

- 1Movie Facts and Fibs (MF$^2$): A Benchmark for Long Movie Understanding葡萄牙里斯本大学高级技术研究所, 葡萄牙电信研究所, 阿姆斯特丹大学ILLC, 阿姆斯特丹大学语言技术实验室, 西班牙国家研究委员会工业机器人与信息学研究所, 北卡罗来纳大学教堂山分校, 哥本哈根大学, 先锋人工智能中心, 赫瑞瓦特大学, 博尔赞-博尔扎诺自由大学, Unbabel, 阿姆斯特丹ELLIS单位, 里斯本ELLIS单位, 特伦托ELLIS单位, 巴塞罗那ELLIS单位 · 2025年



