five

NUBUC

收藏
DataCite Commons2022-05-16 更新2024-07-13 收录
下载链接:
https://catalog.ldc.upenn.edu/LDC2022S04
下载链接
链接失效反馈
官方服务:
资源简介:
<h3>Introduction</h3><br> <p>NUBUC (NyU-BU contextually controlled stories Corpus) was developed by <a href="https://www.nyu.edu/">New York University</a>, <a href="https://www.mpg.de/6971390/empirical-aesthetics">Max Planck Institute for Empirical Aesthetics</a> and <a href="https://www.bu.edu/">Boston University</a>. It contains approximately three hours of English read speech from eight stories focused on linguistic keywords that were created specifically for this corpus, along with transcripts, syntactic annotations and corpus metadata.</p><br> <h3>Data</h3><br> <p>Stories are centered on a protagonist and bear a similarity to a modern fairy tale. Each story consists of approximately 2,000 words organized around critical keywords matched along multiple linguistic dimensions. The story texts comprise a total of 1024 sentences and 16,472 words. Sentences across the eight stories have the same number of words, and alternating sentences contain the linguistically equated keyword. Contextual variables are systematically manipulated while holding linguistic and semantic variables constant across sentences and stories.</p><br> <p>More information about the story design is included in the documentation. The text of the stories, syntactic annotations, and TextGrid word-aligned transcripts are all UTF-8 encoded.</p><br> <p>Each story was read by two different voice actors, one male and one female, in a neutral American English accent. Recordings are 11-12 minutes in duration, for a total of about 90 minutes of continuous speech per speaker. Audio files divided by story as well as by sentence are included in this release; each audio file is presented as a single channel, 11025 Hz, 16-bit, flac compressed wave file.</p><br> <h3>Samples</h3><br> <p>Please view the following samples:</p><br> <ul><br> <li><a href="desc/addenda/LDC2022S04.flac">Speech Sample (FLAC)</a></li><br> <li><a href="desc/addenda/LDC2022S04.txt">Annotation Sample (TXT)</a></li><br> <li><a href="desc/addenda/LDC2022S04.TextGrid">TextGrid Sample</a></li><br> </ul><br> <p>If the speech sample sounds distorted in browser, please download and play locally.</p><br> <h3>Updates</h3><br> <p>None at this time.</p></br> Portions © 2022 MPI Empirical Aesthetics, © 2022 Trustees of the University of Pennsylvania
提供机构:
Linguistic Data Consortium
创建时间:
2022-04-26
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作