遇见数据集

CMU-MOSI、CMU-MOSEI、CMU-MOSEI-OOD

收藏
DataCite Commons2025-12-21 更新2026-05-05 收录
官方服务:

资源简介:

1.The CMU-MOSI dataset is a multimodal dataset created by Carnegie Mellon University for sentiment analysis. Contains 2199 YouTube monologue video clips, covering movie reviews, speeches, and other scenes, with each clip labeled with an emotional rating of [-3, 3]. 2.The CMU-MOSEI dataset is an extended version of the CMU-MOSI dataset. It contains 3228 videos from 1000 speakers, totaling 22856 sentences. 3.The CMU-MOSEI-OOD dataset is constructed using the adaptive simulated annealing algorithm of the CMU-MOSEI dataset, which iteratively adjusts the test distribution to achieve significant differences in word emotion correlation with the training set. This dataset contains two configurations: OOD distribution for second level classification scenarios (labeled CMU-MOSEI *) and OOD distribution for seventh level classification scenarios (labeled CMU-MOSEI * *), corresponding to these two different granularity data partitions. The data annotation is consistent with CMU-MOSEI.

1. CMU-MOSI数据集是由卡内基梅隆大学(Carnegie Mellon University)专为情感分析任务打造的多模态数据集(multimodal dataset)。该数据集收录2199条YouTube平台的独白视频片段,覆盖影评、演讲等多元场景,每条片段均标注有[-3, 3]区间内的情感评级。 2. CMU-MOSEI数据集是CMU-MOSI数据集的扩展版本,其收录了来自1000位发言者的3228条视频,总计22856个句子。 3. CMU-MOSEI-OOD数据集基于CMU-MOSEI数据集的自适应模拟退火算法(adaptive simulated annealing algorithm)构建,该算法通过迭代调整测试分布,使其与训练集在词汇情感相关性层面呈现显著差异。该数据集包含两种配置:适用于二级分类场景的分布外(Out-of-Distribution, OOD)配置(标注为CMU-MOSEI *),以及适用于七级分类场景的分布外配置(标注为CMU-MOSEI * *),二者对应两种不同粒度的数据划分方案,其数据标注规则与CMU-MOSEI数据集完全一致。

提供机构:
Science Data Bank
创建时间:
2025-12-21
搜集汇总
数据集介绍
CMU-MOSI、CMU-MOSEI、CMU-MOSEI-OOD 数据集图片
背景与挑战
背景概述
该系列数据集由CMU-MOSI、CMU-MOSEI和CMU-MOSEI-OOD组成,专注于多模态情感分析。CMU-MOSI包含2199个视频片段,标注情感评分;CMU-MOSEI是其扩展版本,规模更大;CMU-MOSEI-OOD则通过算法构建了与训练集分布显著不同的测试集,用于评估模型在分布外场景下的性能。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务