memebench
收藏资源简介:
MemeBench 是一个用于开放式梗图(Meme)理解的双语诊断基准数据集,旨在系统评估大规模视觉语言模型(LVLM)的文化-语义理解能力。该数据集包含 1,253 个经过专家精细标注的梗图样本,覆盖了动漫/漫画/游戏(ACG)、电影/电视、历史、政治、日常生活、体育和跨领域等 7 个主要文化领域,并同时包含中文和英文内容。数据采用独特的 VIKR 四层结构化标注模式,从四个递进维度对模型理解能力进行诊断:视觉层(Visual,描述所见内容)、身份层(Identity,识别图中实体)、知识层(Knowledge,关联文化背景事实)和推理层(Reasoning,解析幽默机制)。每个样本的标注信息丰富,包括经过人工核实的标准解释(checked_gt)、元数据(如所属领域、语言、梗图类型和结构)、内容安全审核标记,以及四个维度的详细评估检查项和具体内容(如视觉描述与OCR文本、实体身份与来源、文化背景知识、幽默触发点与核心逻辑)。数据集统计显示,在领域分布上,ACG内容占比最高(50.1%),其次是跨领域内容(21.2%);在语言分布上,中文梗图占 61.3%,英文梗图占 38.7%。数据集的创建过程结合了大规模语言模型辅助标注与多领域专家人工验证,并遵循零知识视觉描述、实体双链接等设计原则,确保了标注的一致性和诊断的有效性。该数据集主要用于模型能力的诊断性评估(而非训练),其评估协议采用基于检查表的LLM即法官方法,核心指标为“完全通过率”,要求模型在VIKR四个维度上均正确回答所有检查问题。数据集存在一定的领域、语言和文化偏见,使用时需注意其局限性。
MemeBench is a bilingual diagnostic benchmark dataset for open-ended meme understanding, designed to systematically evaluate the cultural-semantic comprehension capabilities of large vision-language models (LVLMs). This dataset contains 1,253 meme samples meticulously annotated by domain experts, covering 7 major cultural domains including anime/comic/game (ACG), film/television, history, politics, daily life, sports, and cross-domain content, and includes both Chinese and English materials. The dataset employs a unique four-tier structured annotation framework named VIKR, which diagnoses model comprehension capabilities across four progressive dimensions: Visual Layer (describing the visible content), Identity Layer (identifying entities in the image), Knowledge Layer (associating cultural background facts), and Reasoning Layer (analyzing the humor mechanism). Each sample features rich annotation information, including manually verified standard explanations (checked_gt), metadata (such as affiliated domain, language, meme type and structure), content safety audit tags, as well as detailed evaluation check items and specific content across the four dimensions (e.g., visual description and OCR text, entity identity and source, cultural background knowledge, humor trigger points and core logic). Dataset statistics indicate that in terms of domain distribution, ACG content accounts for the largest share (50.1%), followed by cross-domain content (21.2%); in terms of language distribution, Chinese memes make up 61.3% while English memes account for 38.7%. The dataset was constructed through a process combining large language model-assisted annotation and multi-domain expert manual verification, adhering to design principles such as zero-knowledge visual description and entity dual-linking, which ensures annotation consistency and diagnostic validity. This dataset is primarily intended for diagnostic evaluation of model capabilities (rather than model training). Its evaluation protocol adopts the checklist-based LLM-as-judge methodology, with the core metric being "full pass rate", which requires the model to correctly respond to all check questions across all four VIKR dimensions. The dataset exhibits certain domain, language and cultural biases, and users should be mindful of its limitations during usage.





