EXPOTION dataset
收藏资源简介:
EXPOTION数据集是Mohamed bin Zayed University of Artificial Intelligence创建的一个多模态音乐生成数据集,包含7小时同步的视频录音,记录了表情丰富的面部和上半身手势,与相应的音乐对齐。该数据集旨在为未来在多模态和交互式音乐生成领域的研究提供重要支持。数据集由志愿者录制,他们在听30秒音频剪辑时进行面部表情和上半身运动,音频剪辑来自Epidemic Sound的授权乐器曲目。数据集被剪辑成每10秒一个片段,并使用音频字幕模型SALMONN生成每个音频剪辑的字幕,用于训练和推理过程中的文本提示。
The EXPOTION dataset is a multimodal music generation dataset developed by Mohamed bin Zayed University of Artificial Intelligence. It contains 7 hours of synchronized video recordings that capture expressive facial expressions and upper-body gestures, aligned with their corresponding musical accompaniments. This dataset is intended to provide critical support for future research in the fields of multimodal and interactive music generation. The dataset was recorded by volunteer participants, who performed facial expressions and upper-body movements while listening to 30-second audio clips sourced from licensed instrumental tracks provided by Epidemic Sound. The dataset is segmented into 10-second clips, and subtitles for each audio clip are generated using the audio captioning model SALMONN, which are utilized as text prompts for both training and inference workflows.

- 1EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation Mohamed bin Zayed University of Artificial Intelligence, United Arab Emirates · 2025年



