tricky-tts-gemini-flash-tts
收藏资源简介:
该数据集包含文本到语音生成任务的相关数据,主要特征包括:文本提示(text_prompt)、生成的音频数据(generated_audio,采样率为24000Hz)、音频持续时间(duration_s,单位为秒)、音频标记数量(num_audio_tokens)、自动语音识别转录文本(asr_transcription)以及对应的词错误率(asr_wer)和字错误率(asr_cer)。数据集仅包含训练集(train),共4个样本,总大小约4.5MB。数据文件存储路径为data/train-*。该数据集适用于文本到语音合成、语音质量评估等研究任务。
This dataset contains relevant data for text-to-speech generation tasks, with its core features including: text prompt (text_prompt), generated audio data (generated_audio) with a sampling rate of 24000 Hz, audio duration (duration_s, measured in seconds), number of audio tokens (num_audio_tokens), automatic speech recognition (ASR) transcription text (asr_transcription), along with the corresponding word error rate (asr_wer) and character error rate (asr_cer). The dataset only includes the training split (train), totaling 4 samples with an approximate overall size of 4.5 MB. The data files are stored at the path data/train-*. This dataset is suitable for research tasks such as text-to-speech synthesis and speech quality assessment.



