gemini-flash-2.0-speech
收藏资源简介:
Gemini Flash 2.0 Speech数据集是一个高质量的合成语音数据集,由Gemini Flash 2.0模型通过Multimodel Live API生成。该数据集包含两位说话者(Puck和Kore)的英语语音数据,总计47,256个音频文件,总时长为1023527.20秒(约284.31小时)。数据集的文本内容多样,包括数字、LLM生成的句子、维基百科文章、播客、论文、技术和金融新闻以及Reddit帖子。数据集主要用于训练和实验STT/TTS模型。
The Gemini Flash 2.0 Speech Dataset is a high-quality synthetic speech dataset generated by the Gemini Flash 2.0 model via the Multimodal Live API. This dataset contains English speech data from two speakers, Puck and Kore, with a total of 47,256 audio files and a total duration of 1,023,527.20 seconds (approximately 284.31 hours). The dataset includes diverse text content, including numbers, sentences generated by LLMs, Wikipedia articles, podcasts, academic papers, technical and financial news, and Reddit posts. It is primarily used for training and experimenting with STT/TTS models.




