多模态情感聊天翻译数据集(MSCTD)
收藏资源简介:
多模态情感聊天翻译数据集(MSCTD)由北京交通大学和腾讯微信AI模式识别中心共同创建,包含17,841个多模态双语对话,总计173,240个<英语语句, 中文/德语语句, 图像, 情感>四元组。数据集通过自动和人工标注两个步骤构建,确保了数据的质量和多样性。每个语句对都与反映当前对话场景的视觉上下文相对应,并标注有情感标签。MSCTD不仅用于多模态聊天翻译研究,还为多模态对话情感分析提供了新的基准,旨在通过整合对话历史和视觉上下文,生成更准确的翻译,并解决多模态机器翻译在对话中的挑战。
The Multimodal Sentiment Chat Translation Dataset (MSCTD) was co-created by Beijing Jiaotong University and Tencent WeChat AI Pattern Recognition Center. It contains 17,841 multimodal bilingual dialogues, totaling 173,240 <English utterance, Chinese/German utterance, image, sentiment> quadruples. The dataset is constructed through two steps: automatic annotation and manual annotation, to ensure data quality and diversity. Each utterance pair corresponds to a visual context that reflects the current dialogue scenario, and is annotated with sentiment labels. MSCTD not only supports multimodal chat translation research, but also provides a new benchmark for multimodal dialogue sentiment analysis. It aims to generate more accurate translations by integrating dialogue history and visual context, and address the challenges of multimodal machine translation in dialogue scenarios.




