CMU-MOSI、CMU-MOSEI、CMU-MOSEI-OOD
收藏资源简介:
1.The CMU-MOSI dataset is a multimodal dataset created by Carnegie Mellon University for sentiment analysis. Contains 2199 YouTube monologue video clips, covering movie reviews, speeches, and other scenes, with each clip labeled with an emotional rating of [-3, 3]. 2.The CMU-MOSEI dataset is an extended version of the CMU-MOSI dataset. It contains 3228 videos from 1000 speakers, totaling 22856 sentences. 3.The CMU-MOSEI-OOD dataset is constructed using the adaptive simulated annealing algorithm of the CMU-MOSEI dataset, which iteratively adjusts the test distribution to achieve significant differences in word emotion correlation with the training set. This dataset contains two configurations: OOD distribution for second level classification scenarios (labeled CMU-MOSEI *) and OOD distribution for seventh level classification scenarios (labeled CMU-MOSEI * *), corresponding to these two different granularity data partitions. The data annotation is consistent with CMU-MOSEI.
1. CMU-MOSI数据集是由卡内基梅隆大学(Carnegie Mellon University)专为情感分析任务打造的多模态数据集(multimodal dataset)。该数据集收录2199条YouTube平台的独白视频片段,覆盖影评、演讲等多元场景,每条片段均标注有[-3, 3]区间内的情感评级。 2. CMU-MOSEI数据集是CMU-MOSI数据集的扩展版本,其收录了来自1000位发言者的3228条视频,总计22856个句子。 3. CMU-MOSEI-OOD数据集基于CMU-MOSEI数据集的自适应模拟退火算法(adaptive simulated annealing algorithm)构建,该算法通过迭代调整测试分布,使其与训练集在词汇情感相关性层面呈现显著差异。该数据集包含两种配置:适用于二级分类场景的分布外(Out-of-Distribution, OOD)配置(标注为CMU-MOSEI *),以及适用于七级分类场景的分布外配置(标注为CMU-MOSEI * *),二者对应两种不同粒度的数据划分方案,其数据标注规则与CMU-MOSEI数据集完全一致。




