遇见数据集

EmiliaCap-5k

收藏
Zenodo2026-06-12 更新2026-05-26 收录
官方服务:

资源简介:

Description The EmiliaCap-5k dataset provides emotion, pitch, energy, and caption annotations for 2,158,308 audio files from the Emilia-YODAS dataset, corresponding to approximately 5,898 hours of speech. Annotations were generated using a fine-tuned Qwen2-Audio captioning model that leverages chain-of-thought reasoning and a curriculum-learning training schedule. The model was designed to first predict intermediate attributes (emotion, pitch, and energy) before producing the final natural-language caption. The model was trained on roughly 40,000 samples from the TextrolSpeech dataset and demonstrated superior performance compared to 500 alternative configurations, measured using a task-specific caption similarity metric. Contents The dataset consists of a single TSV file with the following fields: path: File identifier following the Emilia-YODAS filename convention emotion: Emotion label (`happy`, `neutral`, `sad`, `angry`, `disgusted`, `surprised`, `contempt`, `fear`) pitch: Pitch label (`low`, `normal`, `high`) energy: Energy label (`low`, `normal`, `high`) caption: Natural-language caption --- Intended Use The EmiliaCap-5k dataset is primarily intended as training and evaluation data for controllable text-to-speech (TTS) systems, where conditioning on emotion, pitch, and energy is required. Other potential use cases include: * Benchmarking audio captioning models that predict fine-grained paralinguistic features* Studying the interaction between emotion, pitch, energy, and natural-language descriptions* Enabling multimodal research that connects speech signals with structured annotations and captions --- Limitations The annotations were automatically generated by a machine learning model and may contain noise or systematic biases. Emotion, pitch, and energy are categorical labels, which may not capture the full nuance of continuous variation. Captions are optimized for similarity-based metrics rather than human readability, and may therefore differ from human-style captions. --- Citation If you use this dataset, please cite it as: @inproceedings{Bountouridis26-BTC, author = {Dimitrios Bountouridis and Filip Packań and Shahin Amiriparian}, title = {{Breaking the Chain: Evaluating CoT and Slot Filling for Speech Captioning and Understanding and Releasing Captions for 5,000 Hours of Speech}}, booktitle = {{Proceedings 34th European Signal Processing Conference (EUSIPCO)}}, year = {2026}, editor = {}, volume = {}, series = {}, pages = {}, address = {Bruges, Belgium}, month = {August}, organization = {EURASIP}, publisher = {IEEE}, note = {}, }

提供机构:
Zenodo
创建时间:
2025-09-17
二维码
社区交流群
二维码
科研交流群
商业服务