MM-SOC
收藏资源简介:
MM-SOC是一个综合性的多模态基准数据集,旨在评估大型语言模型在社交媒体平台上的多模态内容理解能力。该数据集整合了多个著名的多模态数据集,并新增了一个大规模的YouTube标签数据集,涵盖了从错误信息检测、仇恨言论检测到社交上下文生成等多种任务。MM-SOC数据集不仅揭示了当前多模态大型语言模型(MLLMs)的局限性,还为未来模型的改进提供了方向,特别是在提高模型对社会理解能力的需求上。
MM-SOC is a comprehensive multimodal benchmark dataset designed to assess the multimodal content understanding capabilities of large language models on social media platforms. This dataset integrates multiple renowned multimodal datasets and introduces a large-scale YouTube tag dataset, covering a diverse range of tasks including misinformation detection, hate speech detection, and social context generation. The MM-SOC dataset not only reveals the limitations of current multimodal large language models (MLLMs) but also provides valuable directions for future model improvements, especially in addressing the critical need to enhance models' social comprehension abilities.




