Boston University Radio Speech Corpus
收藏资源简介:
<p>The Boston University Radio Speech Corpus was collected primarily to support research in text-to-speech synthesis, particularly generation of prosodic patterns. The corpus consists of professionally read radio news data, including speech and accompanying annotations, suitable for speech and language research. </p><p>The corpus includes speech from seven (four male, three female) FM radio news announcers associated with WBUR, a public radio station. The main radio news portion of the corpus consists of over seven hours of news stories recorded in the WBUR radio studio during broadcasts over a two year period. In addition, the announcers were also recorded in a laboratory at Boston University. In this, the lab news portion, the announcers read a total of 24 stories from the radio news portion. The announcers were first asked to read the stories in their non-radio style and then, 30 minutes later, to read the same stories in their radio style. </p><p>Each story read by an announcer was digitized in paragraph size units, which typically include several sentences. The files were digitized at a 16k Hz sample rate using a 16-bit A/D. The paragraphs were annotated with the orthographic transcription, phonetic alignments, part-of-speech tags and prosodic markers. The orthographic transcripts were generated by hand and include indication of where the speaker took a breath. The phonetic alignments and part-of-speech tags were generated automatically and hand corrected. The prosodic labels were marked by hand and are available only for a subset of the corpus. </p><p>A zipped compressed file <a href="desc/examples/bu_radio_ldc1996s36_sample.ZIP" rel="nofollow">example.zip</a> is available. Please be aware that this file is slightly larger than 1 Mb (1,278,998 bytes). An additional sample file, <a href="./desc/addenda/LDC1996S36.tgz" rel="nofollow">LDC1996.tgz</a> and <a href="./desc/addenda/LDC1996S36.wav" rel="nofollow">WAV sample</a> are also available. </p> </br> Portions © 1996 Trustees of the University of Pennsylvania
波士顿大学广播语音语料库(Boston University Radio Speech Corpus)的采集初衷主要为支持文本到语音合成领域的研究,尤其是韵律模式生成方向的相关工作。该语料库包含专业录制的广播新闻语音数据,配套带有标注信息,适用于语音与语言研究相关场景。 本语料库的语音来源为7名WBUR公共广播电台的新闻播音员(其中男性4名、女性3名)。语料库的核心广播新闻部分,采集自两年播出周期内WBUR广播演播室录制的超过7小时的新闻稿件。此外,研究人员还在波士顿大学的实验室中对这些播音员进行了额外录制:在实验室录制的新闻子语料中,播音员共朗读了来自核心广播新闻部分的24篇稿件,且要求播音员先以非广播播报风格朗读,30分钟后再以标准广播播报风格朗读同一批稿件。 每位播音员朗读的稿件均按段落为单位进行数字化处理,每个段落通常包含多个句子。所有音频文件均以16kHz采样率、16位模数转换(A/D)格式进行数字化。段落标注涵盖正字法转录文本、语音对齐信息、词性标注以及韵律标记。其中正字法转录文本由人工生成,并标注出了播音员换气的位置;语音对齐信息与词性标注由自动生成后经人工校正完成;韵律标签同样由人工手动标注,但仅覆盖语料库的部分子集。 本语料库提供了一个压缩示例文件<example.zip>,文件大小略大于1MB(1,278,998字节)。此外还可获取额外的示例文件:<LDC1996.tgz>与<WAV示例样本>。 部分内容 © 1996 宾夕法尼亚大学托管委员会




