遇见数据集

CSLU: Spoltech Brazilian Portuguese Version 1.0

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

Introduction CSLU: Spoltech Brazilian Portuguese Version 1.0, Linguistic Data Consortium (LDC) catalog number LDC2006S16 and ISBN 1-58563-383-6, contains microphone speech from a variety of regions in Brazil with phonetic and orthographic transcriptions. The utterances consist of both read speech (for phonetic coverage) and responses to questions (for spontaneous speech). The corpus contains 477 speakers and 8,080 separate utterances. A total of 2,540 utterances have been transcribed at the word level (without time alignments), and 5,479 utterances have been transcribed at the phoneme level (with time alignments). Protocol design, recording and transcription were performed by the Universidade Federal do Rio Grande do Sul and the Universidade de Caxias do Sul. Data The data has been recorded at 44.1 kHz (mono, 16-bit) and stored in RIFF format. The recording was conducted with a direct connection from the microphone to the sound card. The sound card was SoundBlaster-compatible. For the prompted sentences, the sentence was hidden from view when recording began, so that the speaker might utter the sentence more naturally. Verification of the recording quality was performed immediately after each utterance recording; the data-collection software allowed the speaker to re-record utterances in case the recording was not of sufficient quality. The acoustic environment was not controlled, in order to allow for background conditions that would occur in application environments. Samples For an example of the data in this corpus, please listen to this audio sample and examine its transcript . Portions © 1994-2002 Center for Spoken Language Understanding, Oregon Health & Science University, © 2006 Trustees of the University of Pennsylvania

引言:CSLU斯波尔特克巴西葡萄牙语语料库V1.0由语言数据联盟(Linguistic Data Consortium,LDC)以目录编号LDC2006S16及ISBN 1-58563-383-6发布,该语料库收录巴西多地区的麦克风采集语音,并附带语音学与正字法转写文本。本语料的话语样本涵盖两类内容:一是用于覆盖语音学标注范围的朗读式语音,二是用于采集自发语音的问答响应语音。该语料库共收录477名说话者的8080条独立话语样本,其中2540条已完成词级转写(未附带时间对齐信息),5479条完成了音素级转写(附带时间对齐信息)。语料的方案设计、录音与转写工作均由南里奥格兰德联邦大学(Universidade Federal do Rio Grande do Sul)与卡希亚斯苏尔大学(Universidade de Caxias do Sul)完成。 数据说明:本语料的语音数据以44.1kHz采样率(单声道、16位量化)录制,并以RIFF格式存储。录音采用麦克风直接连接声卡的方式完成,所用声卡兼容SoundBlaster标准。对于提示语句类样本,录音开始时会隐藏语句文本,以确保说话者能够更自然地完成朗读。每条话语样本录制完成后,会立即进行录音质量验证;若录音质量未达要求,数据采集软件允许说话者重新录制该样本。本次录音未对声学环境进行控制,以模拟实际应用场景中的背景噪声条件。 样本示例:若需查看本语料库的样本示例,请收听对应音频样本并查阅其转写文本。 版权声明:本语料部分内容 © 1994-2002 俄勒冈健康与科学大学口语语言理解中心(Center for Spoken Language Understanding, Oregon Health & Science University),© 2006 宾夕法尼亚大学托管委员会(Trustees of the University of Pennsylvania)。

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务