西班牙语多模态讽刺数据集
收藏资源简介:
本数据集是首个针对西班牙语的多模态讽刺数据集,由赫尔辛基大学数字人文系创建。数据集包含文本、视频和音频,并针对两种西班牙语变体进行了标注,确保了全球语言的广泛方言覆盖。数据集内容包括语音、视频和文本的时间戳,以及每个语音的讽刺标注。创建过程中,使用了JustAnnotate工具进行视频与音频的手动对齐,并修正了原始转录中的错误。该数据集主要应用于西班牙语讽刺检测的研究,旨在通过多模态信息提高讽刺识别的准确性。
This dataset is the first multimodal sarcasm dataset targeting Spanish, developed by the Department of Digital Humanities at the University of Helsinki. The dataset includes text, video, and audio content, and has been annotated for two Spanish varieties to ensure broad dialect coverage of the global Spanish language. It contains timestamps for speech, video, and text, alongside sarcasm annotations for each individual speech utterance. During the dataset construction, the JustAnnotate tool was employed to perform manual alignment between video and audio, and errors within the original transcriptions were rectified. This dataset is primarily utilized for research into Spanish sarcasm detection, with the goal of enhancing the accuracy of sarcasm recognition by leveraging multimodal information.

- 1!Qué maravilla! Multimodal Sarcasm Detection in Spanish: a Dataset and a Baseline赫尔辛基大学数字人文系 · 2021年



