遇见数据集

Emotional speaking style captions for MSP Podcast release 1.12 (EmotionRankCLAP)

收藏
Zenodo2025-09-09 更新2026-05-29 收录
官方服务:

资源简介:

We release a collection of natural-language emotional speaking style descriptions derived from the MSP-Podcast corpus(release 1.12). These descriptions are generated from dimensional emotion attributes (valence and arousal) to help bridge the gap between speech and text modalities in CLAP-style models. Traditionally, speaking style annotations have been limited to speaker traits and categorical emotions. Our dataset extends beyond such fixed categories, enabling the development of emotion understanding models that capture the subtleties of emotion on a continuous scale. All captions were generated using OpenAI’s o1 large language model with the following prompt: “Given the following scale of emotions – valence (1 = very negative; 7 = very positive), arousal (1 = very calm; 7 = very active), write a sentence describing a speaking style that is {VALENCE} on valence and {AROUSAL} on arousal. Do not use any numbers in the sentence. The sentence should start with: The person is speaking ...” Please cite our work : Chandra, S.S., Goncalves, L., Lu, J., Busso, C., Sisman, B. (2025) EmotionRankCLAP: Bridging Natural Language Speaking Styles and Ordinal Speech Emotion via Rank-N-Contrast. Proc. Interspeech 2025, 3000-3004, doi: 10.21437/Interspeech.2025-1198Paper link : https://www.isca-archive.org/interspeech_2025/chandra25_interspeech.pdf MSP Podcast : Lotfian, Reza, and Carlos Busso. "Building naturalistic emotionally balanced speech corpus by retrieving emotional speech from existing podcast recordings." IEEE Transactions on Affective Computing 10.4 (2017): 471-483.

提供机构:
Zenodo
创建时间:
2025-09-08
二维码
社区交流群
二维码
科研交流群
商业服务