遇见数据集

ADEPT: A Dataset for Evaluating Prosody Transfer

收藏
Zenodo2021-07-29 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The ADEPT dataset consists of prosodically-varied natural speech samples for evaluating prosody transfer in english text-to-speech models. The samples include global variations reflecting emotion and interpersonal attitude, and local variations reflecting topical emphasis, propositional attitude, syntactic phrasing and marked tonicity. Txt and wav files are organised according to the folder structure {speech_class}/{subcategory_or_interpretation}/{filename}, where filename follows the naming convention {speaker}_{utterance_id}. Speakers comprise 'ad00' (female voice) and 'ad01' (male voice). For classes with multiple interpretations, we provide the interpretations used in the disambiguation tasks in 'adept_prompts.json'. The corpus only includes prosodic variations that listeners are able to distinguish with reasonable accuracy, and we report these figures as a benchmark against which text-to-speech prosody transfer can be compared. More details can be found in our pre-print about the dataset (https://arxiv.org/abs/2106.08321).

提供机构:
Zenodo
创建时间:
2021-07-20
二维码
社区交流群
二维码
科研交流群
商业服务