DocMSU
收藏资源简介:
DocMSU是一个针对文档级多模态讽刺理解的综合基准数据集,由北京邮电大学创建。该数据集包含102,588条新闻,每条新闻包含文本和图像对,覆盖9个多样化的主题,如健康、商业等。数据集通过爬取知名新闻网站如‘纽约时报’和‘联合国新闻’收集,经过三轮人工标注,确保高质量的标注。DocMSU旨在解决新闻领域中讽刺理解的挑战,特别是在长文本中捕捉讽刺线索的问题,适用于情感分析、假新闻检测和公众舆论分析等领域。
DocMSU is a comprehensive benchmark dataset for document-level multimodal sarcasm understanding, created by Beijing University of Posts and Telecommunications. It contains 102,588 news articles, each paired with a text-image pair, covering 9 diverse topics such as health, business and others. The dataset was collected by crawling well-known news websites including The New York Times and UN News, and underwent three rounds of manual annotation to ensure high-quality labeling. DocMSU aims to address the challenges of sarcasm understanding in the news domain, particularly the difficulty of capturing sarcasm cues in long texts, and is applicable to fields including sentiment analysis, fake news detection and public opinion analysis.




