parlertts_pony_speech_phonemized
收藏资源简介:
该数据集包含多个与语音相关的特征,如说话者信息、文本内容、语音特征等。具体特征包括说话者姓名、来源、开始和结束时间、语音风格、噪音类型、文本内容、持续时间、音高均值和标准差、信噪比、C50、说话速率、音素、STOI、SI-SDR、PESQ、说话者ID、性别、音高、混响、语音单调性、噪音的SDR、语音质量的PESQ、文本描述和音素化文本。数据集分为训练集,包含63947个样本,总大小为40329009字节。
This dataset includes multiple speech-related features, such as speaker information, text content, and speech features. The specific features include speaker name, source, start and end timestamps, speech style, noise type, text content, duration, pitch mean and standard deviation, signal-to-noise ratio, C50, speaking rate, phonemes, STOI, SI-SDR, PESQ, speaker ID, gender, pitch, reverberation, speech monotonicity, SDR of noise, PESQ of speech quality, text description, and phonemized text. The dataset is split into a training set, which contains 63,947 samples with a total size of 40,329,009 bytes.
数据集概述
数据集信息
特征
- speaker: 说话者,类型为字符串。
- source: 来源,类型为字符串。
- start: 开始时间,类型为浮点数(float64)。
- end: 结束时间,类型为浮点数(float64)。
- style: 风格,类型为字符串。
- noise: 噪音,类型为字符串。
- text: 文本,类型为字符串。
- duration: 持续时间,类型为浮点数(float64)。
- utterance_pitch_mean: 语音音调均值,类型为浮点数(float32)。
- utterance_pitch_std: 语音音调标准差,类型为浮点数(float32)。
- snr: 信噪比,类型为浮点数(float64)。
- c50: C50值,类型为浮点数(float64)。
- speaking_rate: 语速,类型为字符串。
- phonemes: 音素,类型为字符串。
- stoi: STOI值,类型为浮点数(float64)。
- si-sdr: SI-SDR值,类型为浮点数(float64)。
- pesq: PESQ值,类型为浮点数(float64)。
- speaker_id: 说话者ID,类型为整数(int32)。
- gender: 性别,类型为字符串。
- pitch: 音调,类型为字符串。
- reverberation: 混响,类型为字符串。
- speech_monotony: 语音单调性,类型为字符串。
- sdr_noise: SDR噪音,类型为字符串。
- pesq_speech_quality: PESQ语音质量,类型为字符串。
- text_description: 文本描述,类型为字符串。
- phonemized_text: 音素化文本,类型为字符串。
数据集分割
- train: 训练集,包含63947个样本,总大小为40329009字节。
数据集大小
- 下载大小: 15351428字节
- 数据集大小: 40329009字节
配置
- config_name: default
- data_files:
- split: train
- path: data/train-*
- data_files:




