nilc-nlp/nurc-sp-prosodic-segmentation
收藏资源简介:
NURC-SP韵律分割数据集是一个巴西葡萄牙语音频数据集,专门用于训练韵律分割模型。该数据集基于NURC-SP项目中的CATNA和Minimal Corpus子语料库进行适配,其中韵律边界使用!!!!!标记。数据集包含三个子集:训练集(train)、清洗集(clean)和测试集(test)。训练集和测试集共同涵盖了CATNA和Minimal Corpus子语料库的全部内容;清洗集则是从训练集中移除了极度嘈杂的音频后生成的。测试集设计为男女说话者数量相等,且一半音频为干净音频,另一半含有中等程度的噪声。
The NURC-SP Prosodic Segmentation Dataset is a Brazilian Portuguese audio dataset specifically tailored for training prosodic segmentation models. It is adapted from the CATNA and Minimal Corpus sub-corpora within the NURC-SP project, where prosodic boundaries are marked with !!!!!. The dataset consists of three subsets: the training set (train), the clean set, and the test set (test). The training and test sets collectively cover the entire content of the CATNA and Minimal Corpus sub-corpora; the clean set is generated by removing extremely noisy audio from the training set. The test set is designed with an equal number of male and female speakers, and half of its audio clips are clean while the other half contain moderate levels of noise.




