NUBUC
收藏资源简介:
<h3>Introduction</h3><br> <p>NUBUC (NyU-BU contextually controlled stories Corpus) was developed by <a href="https://www.nyu.edu/">New York University</a>, <a href="https://www.mpg.de/6971390/empirical-aesthetics">Max Planck Institute for Empirical Aesthetics</a> and <a href="https://www.bu.edu/">Boston University</a>. It contains approximately three hours of English read speech from eight stories focused on linguistic keywords that were created specifically for this corpus, along with transcripts, syntactic annotations and corpus metadata.</p><br> <h3>Data</h3><br> <p>Stories are centered on a protagonist and bear a similarity to a modern fairy tale. Each story consists of approximately 2,000 words organized around critical keywords matched along multiple linguistic dimensions. The story texts comprise a total of 1024 sentences and 16,472 words. Sentences across the eight stories have the same number of words, and alternating sentences contain the linguistically equated keyword. Contextual variables are systematically manipulated while holding linguistic and semantic variables constant across sentences and stories.</p><br> <p>More information about the story design is included in the documentation. The text of the stories, syntactic annotations, and TextGrid word-aligned transcripts are all UTF-8 encoded.</p><br> <p>Each story was read by two different voice actors, one male and one female, in a neutral American English accent. Recordings are 11-12 minutes in duration, for a total of about 90 minutes of continuous speech per speaker. Audio files divided by story as well as by sentence are included in this release; each audio file is presented as a single channel, 11025 Hz, 16-bit, flac compressed wave file.</p><br> <h3>Samples</h3><br> <p>Please view the following samples:</p><br> <ul><br> <li><a href="desc/addenda/LDC2022S04.flac">Speech Sample (FLAC)</a></li><br> <li><a href="desc/addenda/LDC2022S04.txt">Annotation Sample (TXT)</a></li><br> <li><a href="desc/addenda/LDC2022S04.TextGrid">TextGrid Sample</a></li><br> </ul><br> <p>If the speech sample sounds distorted in browser, please download and play locally.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2022 MPI Empirical Aesthetics, © 2022 Trustees of the University of Pennsylvania
<h3>引言</h3><br><p>NUBUC(NyU-BU语境可控故事语料库,NyU-BU contextually controlled stories Corpus)由<a href="https://www.nyu.edu/">纽约大学(New York University)</a>、<a href="https://www.mpg.de/6971390/empirical-aesthetics">马克斯·普朗克经验美学研究所(Max Planck Institute for Empirical Aesthetics)</a>以及<a href="https://www.bu.edu/">波士顿大学(Boston University)</a>联合研发。该语料库包含约3小时的英语朗读语音数据,源自8篇专为该语料库创作、围绕语言关键词展开的故事,同时附带转写文本、句法标注以及语料库元数据。</p><br><h3>数据</h3><br><p>故事均以主人公为核心,风格近似现代童话。每篇故事约含2000词,围绕多个语言维度匹配的关键关键词组织而成。所有故事文本总计包含1024个句子与16472个单词。8篇故事中的句子字数均保持一致,且交替出现的句子包含语言等效的关键词。研究人员通过系统操控语境变量,同时确保不同句子与故事间的语言、语义变量保持恒定。</p><br><p>更多关于故事设计的细节可查阅配套文档。故事文本、句法标注以及与单词对齐的TextGrid转写文本均采用UTF-8编码。</p><br><p>每篇故事均由两名不同的配音演员朗读,分别为男性与女性,采用中性美式英语口音。单篇故事的录音时长约为11至12分钟,每位配音演员的连续语音总时长约为90分钟。本次发布的音频文件按故事及句子进行拆分,所有音频均为单声道、采样率11025Hz、16位精度的FLAC压缩波形音频文件。</p><br><h3>示例</h3><br><p>请查看以下示例:</p><br><ul><br><li><a href="desc/addenda/LDC2022S04.flac">语音示例(FLAC)</a></li><br><li><a href="desc/addenda/LDC2022S04.txt">标注示例(TXT)</a></li><br><li><a href="desc/addenda/LDC2022S04.TextGrid">TextGrid示例</a></li><br></ul><br><p>若浏览器中播放语音示例出现失真,请下载至本地后播放。</p><br><h3>更新记录</h3><br><p>暂无更新记录。</p><br>部分内容 © 2022 马克斯·普朗克经验美学研究所(MPI Empirical Aesthetics),© 2022 宾夕法尼亚大学董事会(Trustees of the University of Pennsylvania)




