MultiMM
收藏资源简介:
MultiMM是一个跨文化多模态隐喻数据集,旨在研究中国和英语中的隐喻。该数据集包含8,461个文本-图像广告对,每个广告对都有精细的注释,提供了对单一文化领域之外的多模态隐喻的深入理解。数据集的创建涉及从商业和公共服务广告中收集文本和视觉元素,并对收集到的数据进行清洗和标注。MultiMM数据集的设计旨在解决自动隐喻处理中的文化偏差问题,并为跨文化多模态隐喻理解提供基准。
MultiMM is a cross-cultural multimodal metaphor dataset designed for studying metaphors in both Chinese and English. This dataset contains 8,461 text-image advertisement pairs, each equipped with fine-grained annotations to deliver in-depth insights into multimodal metaphors beyond single cultural domains. The development of the MultiMM dataset entails collecting textual and visual elements from commercial and public service advertisements, followed by data cleaning and annotation work. The MultiMM dataset is constructed to address cultural bias issues in automatic metaphor processing and serve as a benchmark for cross-cultural multimodal metaphor understanding.
数据集概述
基本信息
- 数据集名称: Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors
- 发布会议: ACL 2025
- 作者: Senqi Yang, Dongyu Zhang, Jing Ren, Ziqi Xu, Xiuzhen Zhang, Yiliao Song, Hongfei Lin, Feng Xia
- 数据集地址: https://github.com/DUTIR-YSQ/MultiMM
数据集描述
- 目的: 用于跨文化多模态隐喻识别和分析,涵盖中文和英文样本。
- 语言: 中文(CN)和英文(EN)
- 数据总量: 8,461条(中文4,397条,英文4,064条)
- 隐喻样本: 4,772条(中文2,583条,英文2,189条)
- 字面样本: 3,689条(中文1,814条,英文1,875条)
- 文本统计:
- 总词数: 213,501(中文145,312,英文68,189)
- 平均词数: 24(中文33,英文15)
- 数据集划分:
- 训练集: 6,768条(中文3,517条,英文3,251条)
- 验证集: 846条(中文440条,英文406条)
- 测试集: 847条(中文440条,英文407条)
数据样本字段
- 图像
- 文本
- 隐喻标签:
1表示隐喻,0表示字面 - 目标域
- 源域
- 情感类型:
1=正面,0=中性,-1=负面
文件结构
code_metaphor/: 包含隐喻检测任务的代码,运行main.py。data/: 包含训练、验证和测试数据。imgs_CN/和imgs_EN/: 图像数据。all/: 原始未分割数据。
相关研究
- MultiMET (ACL 2021): 多模态隐喻理解基础数据集。
- EmoMeta (WWW 2025): 中文多模态隐喻情感分类数据集。
- MultiCMET (EMNLP 2023): 中文多模态隐喻理解基准数据集。
引用格式
bibtex @inproceedings{yang2025cultural, title = {Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors}, author = {Yang, Senqi and Zhang, Dongyu and Ren, Jing and Xu, Ziqi and Zhang, Xiuzhen and Song, Yiliao and Lin, Hongfei and Xia, Feng}, booktitle = {Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)}, pages = {XX--XX}, year = {2025}, address = {Vienna, Austria}, publisher = {Association for Computational Linguistics} }

- 1Cultural Bias Matters: A Cross-Cultural Benchmark Dataset and Sentiment-Enriched Model for Understanding Multimodal Metaphors大连理工大学, 中国; RMIT大学, 澳大利亚; 阿德莱德大学, 澳大利亚 · 2025年



