C2KD
收藏官方服务:
资源简介:
该数据集由三个多模态数据集组成:CREMA-D(情感识别音频-视频数据集)、AVE(视听事件数据集)和 VGGSound(大规模视频-音频数据集)。原始数据为视频格式,需预处理为RGB图像帧和音频波形文件。这些数据集适用于跨模态知识蒸馏任务,旨在通过模态间知识迁移提升模型性能。数据预处理代码位于项目目录的utils/data/下。
This dataset comprises three multimodal datasets: CREMA-D (Audio-Visual Dataset for Emotion Recognition), AVE (Audio-Visual Event Dataset), and VGGSound (Large-Scale Video-Audio Dataset). The raw data is in video format and requires preprocessing into RGB image frames and audio waveform files. These datasets are applicable to cross-modal knowledge distillation tasks, aiming to improve model performance via inter-modal knowledge transfer. The data preprocessing code is located under the utils/data/ directory of the project.
创建时间:
2026-03-22



