MOD
收藏资源简介:
MOD数据集是由中国科学院计算技术研究所智能信息处理重点实验室创建的一个大规模中文多模态对话数据集,专注于将互联网模因融入开放领域对话中,以增强对话的表达力和趣味性。该数据集包含约45,000个对话,总计约606,000条语句,每个对话平均包含13条语句和4个互联网模因,每个含模因的语句都标注了相应的情感。数据集的创建过程涉及从互联网收集模因,并通过专业注释者进行筛选和标注,确保数据质量。该数据集主要用于研究多模态对话建模和情感分析,旨在解决如何使对话系统更加生动和情感丰富的问题。
The MOD dataset is a large-scale Chinese multimodal dialogue dataset created by the Key Laboratory of Intelligent Information Processing, Institute of Computing Technology, Chinese Academy of Sciences. It focuses on integrating internet memes into open-domain dialogues to enhance the expressive power and engagement of conversations. This dataset contains approximately 45,000 dialogues, totaling around 606,000 utterances. On average, each dialogue includes 13 utterances and 4 internet memes, and every utterance with a meme is annotated with its corresponding sentiment. The dataset construction process involves collecting memes from the internet, followed by screening and annotation by professional annotators to ensure data quality. Primarily used for research on multimodal dialogue modeling and sentiment analysis, this dataset aims to solve the problem of how to make dialogue systems more vivid and emotionally rich.

- 1Towards Expressive Communication with Internet Memes: A New Multimodal Conversation Dataset and Benchmark中国科学院计算技术研究所智能信息处理重点实验室 · 2021年



