Ressources for End-to-End French Text-to-Speech Blizzard challenge
收藏资源简介:
Here are 289 chapters of 5 audiobooks from Librivox read by Nadine Eckert-Boulet (NEB): Madame Bovary (MB) by Gustave Flaubert (FL) - 3 volumes, 35 chapters<br> (original wavs; text) Les mystères de Paris (LMP) by Eugene Sue (ES) - 4 volumes, 83 chapters (original wavs1, wavs2, wavs3; text1, text2, text3) Les tribulations d'un chinois en Chine (TCC) by Jules Verne (JV) - 1 volume, 22 chapters (original wavs; text) La fille du pirate (LFDP) by Henri Émile Chevalier (EC) - 7 volumes, 121 chapters (original wavs, text) La vampire (VAMP) by Paul Féval (PF) - 1 volume, 28 chapters (original wavs, text) Each .wav file (sampled at 22050Hz) corresponds to one entire chapter. The format of the filenames is:<br> {author's acronym}_{book's acronym}_{reader's acronym}_{volume's number}_{chapter's number} The NEB_train.csv file gives text and phonetic alignments (essentially for MB and LMP) for utterances in 4 fields separated by '|':<br> {filename}|{start_ms}|{end_ms}|{text or phonetic content}. Most utterances are separated by at least a pause of 400ms. The intervals [start_ms:end_ms] comprise leading and trailing silences of 130ms (since wavs are entire chapters, these silences are "true" ambient silences). When phonetic alignment has been performed, 2 additional fields have been added: {aligned phones}|{durations in ms}. Each input character or phone has a corresponding aligned phone and a duration. Note that all aligned utterances start and end with an aligned phone of 130ms. The set of aligned phones comprises: The set of input phones The silence: '__' The symbol '_' for silent characters, e.g. "chat" is aligned with 's^ _ a _' 29 combined aligned phones ('a&i', 'a&j', 'b&q', 'd&q','d&z', 'd&z^', 'f&q', 'g&q', 'g&z', 'j&i', 'j&u', 'j&q', 'i&j', 'k&q', 'k&s', 'k&s&q', 'l&q', 'm&q', 'n&q', 'r&w', 'r&q', 's&q', 't&q', 't&s', 't&s^', 'w&a', 'z&q', 'p&q') that align to only one character, e.g. "expatrier" is aligned with 'e^ k&s p a t r i&j e _' Text is in UTF8. '«»','¬', '~','""','()','[]' are respectively used for speaking quotes, turn switches, three dots, quoted expression, aside quotes, notes. Because of rare occurrences, 'ö' has been transcribed as 'oe'. Paragraphs (two consecutive carriage returns in the original text) are cued by a special character '§'. It usually ends an utterance but could be used within an utterance if its associated pause is too short. When available, phonetic content is given per word in curly brackets '{}'. We use 39 phonetic symbols: <strong>oral vowels</strong>: a (f<strong><em>a</em></strong>), e (f<em><strong>ée</strong></em>), e^ (f<em><strong>ait</strong></em>), x (f<em><strong>eu</strong></em>), x^ (c<em><strong>oeu</strong></em>r), i (r<em><strong>iz</strong></em>), y (f<em><strong>ut</strong></em>), u (f<em><strong>ou</strong></em>), o (f<em><strong>aux</strong></em>), o^ (p<strong><em>o</em></strong>rc) <strong>schwa</strong>: q (gag<strong><em>e</em></strong>) <strong>nasal vowels</strong>: a~ (r<strong><em>an</em></strong>g), e~ (f<em><strong>in</strong></em>), x~ (<strong><em>un</em></strong>), o~ (r<em><strong>on</strong></em>d) <strong>semi-vowels</strong>: h (h<em><strong>u</strong></em>it), w (<strong><em>ou</em></strong>ate), j (h<em><strong>i</strong></em>er) <strong>consonants</strong>: p (<em><strong>p</strong></em>as), t (<strong><em>t</em></strong>as), k (<em><strong>c</strong></em>as), b (<strong><em>b</em></strong>as), d (<em><strong>d</strong></em>os), g (<em><strong>g</strong></em>ars), f (<em><strong>f</strong></em>aux), s (<strong><em>s</em></strong>ot) , s^ (<strong><em>ch</em></strong>at), v (<strong><em>v</em></strong>u), z (<strong><em>z</em></strong>ut), z^ (<em><strong>j</strong></em>us), r (<strong><em>r</em></strong>iz), l (<em><strong>l</strong></em>a), m (<strong><em>m</em></strong>a), n (<strong><em>n</em></strong>on), n~ (oi<strong><em>gn</em></strong>on), ng (campi<em><strong>ng</strong></em>)
本数据集包含来自Librivox平台、由纳丁·埃克尔特-布勒(Nadine Eckert-Boulet,缩写NEb)录制的5部法语有声书,共计289个章节,具体信息如下: 1. 《包法利夫人》(Madame Bovary,缩写MB),作者居斯塔夫·福楼拜(Gustave Flaubert,缩写FL):共3卷,35个章节,包含原始WAV音频文件与对应文本文件; 2. 《巴黎的秘密》(Les mystères de Paris,缩写LMP),作者欧仁·苏(Eugene Sue,缩写ES):共4卷,83个章节,包含三组原始WAV音频文件(wavs1、wavs2、wavs3)与三组对应文本文件(text1、text2、text3); 3. 《旅华中国男子历险记》(Les tribulations d'un chinois en Chine,缩写TCC),作者儒勒·凡尔纳(Jules Verne,缩写JV):共1卷,22个章节,包含原始WAV音频文件与对应文本文件; 4. 《海盗的女儿》(La fille du pirate,缩写LFDP),作者亨利·埃米尔·谢瓦利埃(Henri Émile Chevalier,缩写EC):共7卷,121个章节,包含原始WAV音频文件与对应文本文件; 5. 《吸血鬼》(La vampire,缩写VAMP),作者保罗·费瓦尔(Paul Féval,缩写PF):共1卷,28个章节,包含原始WAV音频文件与对应文本文件。 每个.wav音频文件采样率为22050Hz,对应完整的一个章节。文件名格式为:{作者缩写}_{书籍缩写}_读者缩写_{卷号}_{章节号}。 文件NEb_train.csv以竖线「|」分隔的4个字段,提供了语音片段的文本与语音对齐信息(主要覆盖《包法利夫人》与《巴黎的秘密》),各字段依次为:{文件名}|{起始毫秒数}|{结束毫秒数}|{文本或语音内容}。绝大多数语音片段之间的间隔至少为400ms的停顿。区间[start_ms:end_ms]包含了130ms的首尾静音段——由于单音频文件对应完整章节,此类静音均为真实环境背景音。 当完成语音对齐处理后,文件会额外增加两个字段:{对齐音素}|{毫秒时长}。每个输入字符或音素均对应一个对齐音素与对应的时长。需注意,所有对齐后的语音片段均以时长130ms的静音音素作为开头与结尾。 对齐音素集合包含以下三类: 1. 基础输入音素集合; 2. 静音音素:`__`; 3. 用于表示静默字符的符号`_`,例如单词「chat」会被对齐为`s^ _ a _`。 此外还包含29个组合对齐音素:`a&i`、`a&j`、`b&q`、`d&q`、`d&z`、`d&z^`、`f&q`、`g&q`、`g&z`、`j&i`、`j&u`、`j&q`、`i&j`、`k&q`、`k&s`、`k&s&q`、`l&q`、`m&q`、`n&q`、`r&w`、`r&q`、`s&q`、`t&q`、`t&s`、`t&s^`、`w&a`、`z&q`、`p&q`,此类组合音素仅对应单个字符,例如单词「expatrier」会被对齐为`e^ k&s p a t r i&j e _`。 文本采用UTF8编码。其中符号「«»」、「¬」、「~」、`""`、`()`、`[]`分别用于标记说话引号、话轮切换、省略号、直接引语、旁注引号与注释。在稀有场景下,字符「ö」会被转写为「oe」。 原文中的段落(即原文本中连续两个回车换行)由特殊字符「§」标记。该符号通常作为一个语音片段的结尾,但若关联的停顿过短,也可出现在语音片段内部。 若提供语音内容,则会按单词用花括号`{}`括起。本数据集共使用39个音素符号,分类如下: - **口腔元音**:a(如fat的/a/)、e(如fée的/e/)、e^(如ait的/ɛ/)、x(如feu的/ø/)、x^(如cœur的/œ/)、i(如riz的/i/)、y(如fut的/y/)、u(如fou的/u/)、o(如aux的/o/)、o^(如porc的/ɔ/); - **央元音(Schwa)**:q(如gag的/ə/); - **鼻元音**:a~(如rang的/ɑ̃/)、e~(如fin的/ɛ̃/)、x~(如un的/œ̃/)、o~(如rond的/ɔ̃/); - **半元音**:h(如huit的/ɥ/)、w(如ouate的/w/)、j(如hier的/j/); - **辅音**:p(如pas的/p/)、t(如tas的/t/)、k(如cas的/k/)、b(如bas的/b/)、d(如dos的/d/)、g(如gars的/ɡ/)、f(如faux的/f/)、s(如sot的/s/)、s^(如chat的/ʃ/)、v(如vu的/v/)、z(如zut的/z/)、z^(如jus的/ʒ/)、r(如riz的/ʁ/)、l(如la的/l/)、m(如ma的/m/)、n(如non的/n/)、n~(如oignon的/ɲ/)、ng(如camping的/ŋ/)。



