foutput
收藏资源简介:
该数据集是一个多语言、多说话人音频数据集,包含49个独立配置,每个配置对应一种特定语言和说话人组合。语言涵盖德语(de)、西班牙语(es)、法语(fr)、俄语(ru)、土耳其语(tu)以及代码为pu的语言(具体语种未明确)。每个配置下包含一位说话人的音频数据,说话人以s1、s2等编号标识。所有配置仅提供训练集,样本数量在不同配置间差异较大,从436个(fr_s5)到3269个(tu_s6)不等,总样本数可通过累加各配置样本数获得。数据以分片文件形式存储,路径格式为语言代码/说话人编号/train-*。每个样本包含两个字段:audio字段为采样率24kHz的音频数据;folder字段为字符串类型,可能表示原始文件目录信息。该数据集适用于多语言语音处理、说话人识别、语音合成等任务,但缺乏对音频内容、录制条件、文本转录等元数据的详细描述。
This dataset is a multilingual, multi-speaker audio dataset consisting of 49 independent configurations, each corresponding to a specific language-speaker combination. Languages covered include German (de), Spanish (es), French (fr), Russian (ru), Turkish (tu), and a language with the code pu (specific language not specified). Each configuration contains audio data from one speaker, with speakers identified by identifiers such as s1, s2, etc. All configurations only provide training splits. The number of samples varies significantly across configurations, ranging from 436 (fr_s5) to 3269 (tu_s6). The total number of samples can be obtained by summing the sample counts of each configuration. The data is stored as sharded files, with the path format being language_code/speaker_id/train-*. Each sample contains two fields: the `audio` field holds audio data with a sampling rate of 24 kHz; the `folder` field is a string type, which may represent the original file directory information. This dataset is applicable to tasks such as multilingual speech processing, speaker recognition, and speech synthesis, but lacks detailed descriptions of metadata such as audio content, recording conditions, and text transcriptions.




