MUStARD++
收藏资源简介:
MUStARD++是一个多模态讽刺检测数据集,由萨里大学创建,旨在通过语言、语音和视觉线索全面捕捉讽刺现象。数据集包含1202个视频样本,来源于多个流行电视节目,通过手动标注确保高质量的讽刺标签。创建过程中,研究者们通过多轮标注和验证确保数据的准确性和多样性。该数据集主要应用于自动讽刺检测,帮助机器理解并识别讽刺语境,解决讽刺识别中的多模态挑战。
MUStARD++ is a multimodal sarcasm detection dataset created by the University of Surrey, designed to comprehensively capture sarcasm via linguistic, acoustic and visual cues. The dataset comprises 1202 video samples sourced from multiple popular television programs, with high-quality sarcasm annotations ensured through manual labeling. During its development, researchers conducted multiple rounds of annotation and validation to guarantee the accuracy and diversity of the dataset. This dataset is primarily utilized for automated sarcasm detection, assisting machines in comprehending and identifying sarcastic contexts, and addressing the multimodal challenges in sarcasm recognition.




