遇见数据集

C2KD

收藏
Hugging Face2026-03-22 更新2026-03-23 收录
官方服务:

资源简介:

该数据集由三个多模态数据集组成:CREMA-D(情感识别音频-视频数据集)、AVE(视听事件数据集)和 VGGSound(大规模视频-音频数据集)。原始数据为视频格式,需预处理为RGB图像帧和音频波形文件。这些数据集适用于跨模态知识蒸馏任务,旨在通过模态间知识迁移提升模型性能。数据预处理代码位于项目目录的utils/data/下。

This dataset comprises three multimodal datasets: CREMA-D (Audio-Visual Dataset for Emotion Recognition), AVE (Audio-Visual Event Dataset), and VGGSound (Large-Scale Video-Audio Dataset). The raw data is in video format and requires preprocessing into RGB image frames and audio waveform files. These datasets are applicable to cross-modal knowledge distillation tasks, aiming to improve model performance via inter-modal knowledge transfer. The data preprocessing code is located under the utils/data/ directory of the project.

创建时间:
2026-03-22
二维码
社区交流群
二维码
科研交流群
商业服务