m3eval
收藏资源简介:
M³Eval(多模态记忆评估)是一个专门设计用于系统评估多模态模型记忆能力的视频问答基准数据集。该数据集旨在填补现有视频理解研究在记忆评估方面的空白,重点关注模型在长视频处理中保留了什么信息、信息的保真度以及记忆在干扰下的鲁棒性。其设计基于认知心理学原理,通过一系列精心构建的任务来分离和评估记忆的不同关键维度。数据集包含多种评估任务:1) 分心注意力:要求模型同时记忆并排播放的两个独立视频流中的内容;2) 记忆干扰:探究顺序呈现的视频之间如何产生前摄干扰(先学信息干扰后学信息)或倒摄干扰(后学信息干扰先学信息);3) 交错事件:评估模型从时间上交错混合的多个视频片段中重建原始事件序列的能力;4) N-Back:一种经典的认知任务变体,要求模型判断当前视频片段是否与之前第N个位置的片段匹配,用于测试符号记忆。数据形式为多模态(视频)输入,任务类型主要为视觉问答和多项选择。数据集规模在1,000到10,000个样本之间,适用于对多模态模型(尤其是针对长视频理解的模型)的记忆机制进行深入分析和基准测试。
M³Eval (Multimodal Memory Evaluation) is a video question-answering benchmark dataset specifically designed for systematically evaluating the memory capabilities of multimodal models. The dataset aims to fill the gap in existing video understanding research regarding memory assessment, focusing on what information models retain during long video processing, the fidelity of that information, and the robustness of memory under interference. Its design is based on cognitive psychology principles, using a series of carefully constructed tasks to isolate and evaluate different key dimensions of memory. The dataset includes various evaluation tasks: 1) Distraction Attention: requires models to simultaneously memorize content from two independent video streams played side by side; 2) Memory Interference: explores how sequentially presented videos cause proactive interference (prior learning interfering with subsequent learning) or retroactive interference (subsequent learning interfering with prior learning); 3) Interleaved Events: assesses the ability of models to reconstruct the original event sequence from multiple video segments interleaved temporally; 4) N-Back: a variant of a classic cognitive task, requiring models to judge whether the current video segment matches the segment from N positions earlier, used to test symbolic memory. The data format is multimodal (video) input, with task types primarily being visual question answering and multiple choice. The dataset scale ranges from 1,000 to 10,000 samples, suitable for in-depth analysis and benchmarking of memory mechanisms in multimodal models, especially those targeting long video understanding.




