LDC Spoken Language Sampler - Fifth Release
收藏资源简介:
Introduction LDC (Linguistic Data Consortium) Spoken Language Sampler - Fifth Release contains samples from 19 corpora published by LDC between 1996 and 2019. LDC distributes a wide and growing assortment of resources for researchers, engineers and educators whose work is concerned with human languages. Historically, most linguistic resources were not generally available to interested researchers but were restricted to single laboratories or to a limited number of users. Inspired by the success of selected readily-available and well-known data sets, such as the Brown University text corpus, LDC was founded in 1992 to provide a new mechanism for large-scale corpus development and resource sharing. With the support of its members, LDC provides critical services to the language research community that include: maintaining the LDC data archives, producing and distributing data via media or web download, negotiating intellectual property agreements with potential information providers and maintaining relations with other like-minded groups around the world. Resources available from LDC include speech, text, video and lexicons in multiple languages, as well as software tools to facilitate the use of corpus materials. For a complete view of LDC's publications, browse the Catalog. The sampler is available as a free download. Data The LDC Spoken Language Sampler - Fifth Release provides speech and transcript samples and is designed to illustrate the variety and breadth of the speech-related resources available from the LDC Catalog. The sound files included in this release are excerpts that have been modified in various ways relative to the original data as published by LDC: Most excerpts are truncated to be much shorter than the original files, typically about 2 minutes. Samples shorter than this typically represent the entirety of a single file. Signal amplitude has been adjusted where necessary to normalize playback volume. Some corpora are published in compressed form, but all samples here are uncompressed. Some text files are presented as images to ensure foreign character sets display properly. In the below table, the link for the catalog number takes you to the catalog entry for that corpus. LDC2018S06 2011 NIST Language Recognition Evaluation Test Set 2011 NIST Language Recognition Evaluation Test Set contains selected training data and the evaluation test set for the 2011 NIST Language Recognition Evaluation. It consists of approximately 204 hours of conversational telephone speech and broadcast audio collected by the Linguistic Data Consortium (LDC) in the following 24 languages and dialects: Arabic (Iraqi), Arabic (Levantine), Arabic (Maghrebi), Arabic (Standard), Bengali, Czech, Dari, English (American), English (Indian), Farsi, Hindi, Lao, Mandarin, Punjabi, Pashto, Polish, Russian, Slovak, Spanish, Tamil, Thai, Turkish, Ukrainian and Urdu. LDC2018S14 AISHELL-1 AISHELL-1 contains approximately 520 hours of Chinese Mandarin speech from 400 speakers recorded simultaneously on three different devices with associated transcripts. The goal of the collection was to support speech recognition system development in domains such as smart homes, autonomous driving, entertainment, finance, and science and technology. LDC2018S15 Avatar Education Portuguese Avatar Education Portuguese contains approximately 80 minutes of Brazilian Portuguese microphone speech with phonetic and orthographic transcriptions. The data was developed for Avatar Education, an animated virtual assistant designed to enhance communication and interaction in educational contexts, such as online learning. LDC96S60 CALLFRIEND Vietnamese CALLFRIEND Vietnamese consists of approximately 60 unscripted telephone conversations between native speakers of Vietnamese. The duration of each conversation was between 5-30 minutes. The corpus also includes documentation describing speaker information (sex, age, education, callee telephone number) and call information (channel quality, number of speakers. LDC2019S07 CIEMPIESS Experimentation CIEMPIESS (Corpus de Investigación en Español de México del Posgrado de Ingeniería Eléctrica y Servicio Social) Experimentation was developed at the National Autonomous University of Mexico (UNAM) and consists of approximately 22 hours of Mexican Spanish broadcast and read speech with associated transcripts. The goal of this work was to create acoustic models for automatic speech recognition. LDC97S63 The CMU Kids Corpus The CMU Kids Corpus was developed in 1995-1996 and is a database of sentences read aloud by 76 children, totaling 5,180 utterances. This data set was designed as a training set of children's speech for the SPHINX II automatic speech recognizer in the LISTEN project at Carnegie Mellon University. LDC2008S01 CSLU: Portland Cellular Telephone Speech Version 1.3 Created by the Center for Spoken Language Understanding (CSLU) at Oregon Health and Science University, CSLU: Portland Cellular Telephone Speech Version 1.3 is a collection of cellular telephone speech (7,571 utterances) and corresponding orthographic and phonetic transcriptions. LDC2018S01 DIRHA English WSJ Audio DIRHA English WSJ Audio is comprised of approximately 85 hours of real and simulated read speech by six native American English speakers. It was developed as part of the Distant-Speech Interaction for Robust Home Applications (DIRHA) Project, which addressed natural spontaneous speech interaction with distant microphones in a domestic environment. LDC2019S14 The DKU-JNU-EMA Electromagnetic Articulography Database The DKU-JNU-EMA Electromagnetic Articulography Database was developed by Duke Kunshan University and Jinan University and contains approximately 10 hours of articulography and speech data in Mandarin, Cantonese, Hakka, and Teochew Chinese from two to seven native speakers for each dialect. LDC2002S28 Emotional Prosody Speech and Transcripts Emotional Prosody Speech and Transcripts was developed by LDC and contains audio recordings and corresponding transcripts, designed to support research in emotional prosody and collected over an eight-month period in 2000-2001. The recordings consist of professional actors reading a series of semantically neutral utterances (dates and numbers) spanning 14 distinct emotional categories. LDC2019S09 First DIHARD Challenge Development - Eight Sources First DIHARD Challenge Development - Eight Sources was developed by LDC and contains approximately 17 hours of English and Chinese speech data along with corresponding annotations used in support of the First DIHARD Challenge. This release, when combined with First DIHARD Challenge Development - SEEDLingS (LDC2019S10), contains the development set audio data and annotation (diarization, segmentation) as well as the official scoring tool. LDC2017S19 IARPA Babel Zulu Language Pack IARPA-babel206b-v0.1e IARPA Babel Zulu Language Pack IARPA-babel206b-v0.1e was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 211 hours of Zulu conversational and scripted telephone speech collected in 2012 and 2013 along with corresponding transcripts. LDC2004S02 ICSI Meeting Speech ICSI Meeting Speech contains approximately 72 hours of speech from 53 unique speakers in 75 meetings collected at Berkeley’s International Computer Science Institute (ICSI) in 2000-2002. The recordings were made during regular weekly meetings of various ICSI working teams, including the team working on the ICSI Meeting Project. The speech files range in length from 17 to 103 minutes, but in general are less than one hour each. LDC2012S04 Malto Speech and Transcripts Malto Speech and Transcripts contains approximately 8 hours of Malto speech data collected between 2005 and 2009 from 27 speakers (22 males, 5 females), accompanying transcripts, English translations and glosses for 6 hours of the collection. Speakers were asked to talk about themselves, their lives, rituals and folklore; elicitation interviews were then conducted. The goal of the work was to present the current state and dialectal variation of Malto. LDC2018S08 Multi-Language Conversational Telephone Speech 2011 -- Central European Multi-Language Conversational Telephone Speech 2011 -- Central European was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 44 hours of telephone speech in two distinct language varieties of Central Europe: Czech and Slovak. The data was collected to support research and technology evaluation in automatic language identification, specifically language pair discrimination for closely related languages/dialects. Portions of these telephone calls were used in the NIST 2011 Language Recognition Evaluation. LDC2006S13 N4 NATO Native and Non-Native Speech N4 NATO Native and Non-Native Speech corpus was developed by the NATO research group on Speech and Language Technology in order to provide a military-oriented database for multilingual and non-native speech processing studies. It consists of 115 native and non-native speakers using NATO English procedure between ships and reading from a text, "The North Wind and the Sun," in both English and the speaker's native language. LDC2018S10 RATS Language Identification RATS Language Identification was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 5,400 hours of Levantine Arabic, Farsi, Dari, Pashto and Urdu conversational telephone speech with annotation of speech segments. The corpus was created to provide training, development and initial test sets for the Language Identification (LID) task in the DARPA RATS (Robust Automatic Transcription of Speech) program. LDC2012S06 Turkish Broadcast News Speech and Transcripts Turkish Broadcast News Speech and Transcripts was developed by Boğaziçi University, Istanbul, Turkey and contains approximately 130 hours of Voice of America (VOA) Turkish radio broadcasts and corresponding transcripts. This is part of a larger corpus of Turkish broadcast news data collected and transcribed with the goal to facilitate research in Turkish automatic speech recognition and its applications. The VOA material was collected between December 2006 and June 2009 using a PC and TV/radio card setup. The data collected during the period 2006-2008 was recorded from analog FM radio; the 2009 broadcasts were recorded from digital satellite transmissions. LDC2017S17 Vehicle City Voices Corpus – Part I Vehicle City Voices Corpus – Part I was developed at the University of Michigan-Flint, and is an ongoing oral history project and survey of English language variation in Flint, Michigan. It contains approximately 16 hours of speech with corresponding transcripts from 21 interviews of Flint residents conducted between 2012 and 2015. The corpus was designed to provide high-quality recordings for acoustic analysis and to examine narrative structure and discursive construction of individual and collective identity in urban spaces. Portions © 2019 Trustees of the University of Pennsylvania
引言 第五版语言数据联盟(Linguistic Data Consortium,LDC)口语语言采样集收录了1996年至2019年间LDC发布的19个语料库的采样样本。 LDC为关注人类语言研究的科研人员、工程师与教育工作者提供了日益丰富的各类资源。过往多数语言资源仅对单一实验室或少量特定用户开放,感兴趣的研究者难以普遍获取。受布朗大学文本语料库等一批易于获取的知名数据集成功案例的启发,LDC于1992年成立,旨在为大规模语料库开发与资源共享提供全新机制。在会员单位的支持下,LDC为语言研究社区提供多项核心服务:维护LDC数据档案库、通过介质分发或网络下载的方式发布数据、与潜在信息提供者洽谈知识产权协议,以及与全球其他志同道合的学术团体保持合作关系。LDC可提供的资源涵盖多语言的语音、文本、视频与词典,以及助力语料材料使用的软件工具。如需完整查看LDC的所有发布资源,请浏览其目录页面。本采样集可免费下载获取。 数据说明 第五版LDC口语语言采样集提供语音与转录文本样本,旨在展示LDC目录中语音相关资源的多样性与覆盖范围。本版包含的音频文件均为经多种方式处理后的原始发布数据节选:大部分节选内容已被截断,时长远短于原始文件,通常约为2分钟;短于此时长的样本一般为单个文件的完整内容。必要时已调整信号幅值以统一播放音量。部分语料库以压缩格式发布,但本采样集中的所有样本均为未压缩格式。部分文本文件以图像形式呈现,以确保外文字符集能够正常显示。下表中,目录编号的链接可跳转至对应语料库的目录详情页。 LDC2018S06 2011年美国国家标准与技术研究院(National Institute of Standards and Technology, NIST)语言识别评估测试集 该数据集为2011年NIST语言识别评估任务提供精选训练数据与评估测试集,包含LDC收集的约204小时会话电话语音与广播音频,覆盖以下24种语言及方言:伊拉克阿拉伯语、黎凡特阿拉伯语、马格里布阿拉伯语、标准阿拉伯语、孟加拉语、捷克语、达里语、美式英语、印度式英语、波斯语、印地语、老挝语、普通话、旁遮普语、普什图语、波兰语、俄语、斯洛伐克语、西班牙语、泰米尔语、泰语、土耳其语、乌克兰语与乌尔都语。 LDC2018S14 AISHELL-1 该数据集包含来自400名说话者的约520小时普通话语音数据,由三种不同设备同步录制,并配有对应转录文本。该数据集的采集旨在支持智能家居、自动驾驶、娱乐、金融与科技等领域的语音识别系统研发。 LDC2018S15 阿凡达教育葡萄牙语数据集 该数据集包含约80分钟巴西葡萄牙语麦克风语音数据,配有语音转写与正字法转录文本。该数据集为阿凡达教育项目开发,该项目是一款动画虚拟助手,旨在提升在线学习等教育场景中的沟通与交互体验。 LDC96S60 CALLFRIEND越南语数据集 该数据集包含约60场越南语母语者之间的无脚本电话对话,每场对话时长为5至30分钟。该语料库还包含说话者信息(性别、年龄、教育程度、被叫电话号码)与通话信息(信道质量、说话者数量)的相关文档。 LDC2019S07 CIEMPIESS实验数据集(全称:Corpus de Investigación en Español de México del Posgrado de Ingeniería Eléctrica y Servicio Social) 该数据集由墨西哥国立自治大学(National Autonomous University of Mexico, UNAM)开发,包含约22小时墨西哥西班牙语广播与朗读语音数据,配有对应转录文本。该数据集的研发目标是为自动语音识别系统构建声学模型。 LDC97S63 CMU儿童语料库 该语料库于1995至1996年间开发,包含76名儿童朗读的共5180条语句的语音数据。该数据集专为卡内基梅隆大学LISTEN项目中的SPHINX II自动语音识别器设计,作为儿童语音训练数据集使用。 LDC2008S01 CSLU:波特兰蜂窝电话语音数据集V1.3 该数据集由俄勒冈健康与科学大学口语语言理解中心(Center for Spoken Language Understanding, CSLU)开发,包含7571条蜂窝电话语音数据及对应的正字法与语音转写文本。 LDC2018S01 DIRHA英语WSJ音频数据集 该数据集包含6名美式英语母语者录制的约85小时真实与模拟朗读语音数据。该数据集作为稳健家居应用远程语音交互(Distant-Speech Interaction for Robust Home Applications, DIRHA)项目的一部分开发,旨在实现家庭环境下远程麦克风的自然自发语音交互。 LDC2019S14 DKU-JNU-EMA电磁发音描记数据库 该数据库由昆山杜克大学与暨南大学联合开发,包含约10小时的发音描记与语音数据,覆盖普通话、粤语、客家话与潮州话四种汉语方言,每种方言由2至7名母语说话者录制。 LDC2002S28 情感韵律语音与转录文本数据集 该数据集由LDC开发,包含2000至2001年间历时8个月采集的音频录音与对应转录文本,旨在支持情感韵律相关研究。该数据集的录音由专业演员朗读一系列语义中性的语句(日期与数字),涵盖14种不同的情感类别。 LDC2019S09 首届DIHARD挑战开发数据集——八源版 该数据集由LDC开发,包含约17小时英语与汉语语音数据及对应标注信息,用于支撑首届DIHARD挑战赛事。本版数据集与首届DIHARD挑战开发数据集——SEEDLingS(LDC2019S10)结合后,可提供开发集音频数据、标注(含说话人diarization与分段标注)以及官方评分工具。 LDC2017S19 IARPA Babel祖鲁语语言包IARPA-babel206b-v0.1e 该语言包由Appen为美国情报高级研究项目活动(Intelligence Advanced Research Projects Activity, IARPA)Babel项目开发,包含2012至2013年间采集的约211小时祖鲁语会话与脚本电话语音数据及对应转录文本。 LDC2004S02 ICSI会议语音数据集 该数据集包含2000至2002年间在伯克利国际计算机科学研究所(International Computer Science Institute, ICSI)采集的75场会议的约72小时语音数据,涉及53名不同说话者。录音来自ICSI各工作组的常规周会,包括ICSI会议项目团队的会议。语音文件时长从17至103分钟不等,总体单文件时长均少于1小时。 LDC2012S04 马尔托语语音与转录文本数据集 该数据集包含2005至2009年间从27名说话者(22名男性、5名女性)处采集的约8小时马尔托语语音数据,配有对应转录文本,其中6小时数据还附带英语翻译与词汇注释。采集过程中,首先要求说话者讲述自身经历、生活、仪式与民俗,随后进行引导式访谈。该数据集旨在展示马尔托语的当前使用现状与方言差异。 LDC2018S08 2011年多语言会话电话语音数据集——中欧版 该数据集由LDC开发,包含约44小时中欧两种不同语言变体的电话语音数据:捷克语与斯洛伐克语。该数据集用于支撑自动语言识别领域的研究与技术评估,尤其是近缘语言/方言的语言对判别任务。部分通话数据曾用于2011年NIST语言识别评估任务。 LDC2006S13 N4北约母语与非母语语音语料库 该语料库由北约语音与语言技术研究小组开发,旨在为多语言与非母语语音处理研究提供面向军事场景的数据库。该数据集包含115名母语与非母语说话者,他们按照北约英语通信流程进行舰间通话,并朗读《北风与太阳》的英文原文与各自母语版本。 LDC2018S10 RATS语言识别数据集 该数据集由LDC开发,包含约5400小时黎凡特阿拉伯语、波斯语、达里语、普什图语与乌尔都语会话电话语音数据,并配有语音片段标注。该语料库旨在为美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)鲁棒语音自动转录(Robust Automatic Transcription of Speech, RATS)项目中的语言识别(Language Identification, LID)任务提供训练集、开发集与初始测试集。 LDC2012S06 土耳其广播新闻语音与转录文本数据集 该数据集由土耳其伊斯坦布尔博加齐奇大学开发,包含约130小时美国之音(Voice of America, VOA)土耳其语广播音频及对应转录文本。该数据集是更大规模土耳其广播新闻语料库的一部分,该语料库的采集与转录旨在推动土耳其语自动语音识别及其应用的研究。VOA素材采集于2006年12月至2009年6月,采用PC与电视/广播卡设备完成录制:2006至2008年采集的数据来自模拟FM广播,2009年的广播数据则来自数字卫星传输。 LDC2017S17 汽车城之声语料库——第一部分 该语料库由密歇根大学弗林特分校开发,是一项正在进行中的口述历史项目与英语语言变异调研项目,调研范围为密歇根州弗林特市的英语使用情况。该数据集包含2012至2015年间对21名弗林特居民的访谈录音及对应转录文本,总时长约16小时。该语料库旨在为声学分析提供高质量录音,并探究城市空间中个体与集体身份的叙事结构与话语建构。部分内容 © 2019 宾夕法尼亚大学托管会




