遇见数据集

USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition

收藏
DataCite Commons2021-07-01 更新2025-04-16 收录
官方服务:

资源简介:

<h3>Introduction</h3><br> <p>USC-SFI MALACH Interviews and Transcripts English &ndash; Speech Recognition Edition, LDC Catalog Number LDC2019S11 and ISBN 1-58563-889-7, was developed by IBM as part of the <a href="http://malach.umiacs.umd.edu/">MALACH (Multilingual Access to Large Spoken ArCHives) Project</a>. This edition augments USC-SFI MALACH Interviews and Transcripts English (<a href="../../../LDC2012S05">LDC2012S05</a>) by modifying and updating a subset of the original corpus for use with the <a href="https://kaldi-asr.org/">Kaldi</a> toolkit in speech recognition work, and is easily portable for use by other speech recognition systems as well. It contains approximately 168 hours of interviews from 682 Holocaust witnesses along with transcripts, a lexicon, Kaldi specific files, and other documentation.</p><br> <p>Inspired by his experience making <em>Schindler&rsquo;s List</em>, Steven Spielberg established the Survivors of the Shoah Visual History Foundation in 1994 to gather video testimonies from survivors and other witnesses of the Holocaust. While most of those who gave testimony were Jewish survivors, the Foundation also interviewed homosexual survivors, Jehovah&rsquo;s Witness survivors, liberators and liberation witnesses, political prisoners, rescuers and aid providers, Roma and Sinti (Gypsy) survivors, survivors of eugenics policies, and war crimes trials participants. The Foundation&rsquo;s Visual History Archive holds nearly 55,000 video testimonies in 43 languages, representing 65 countries; it is the largest archive of its kind in the world. In 2006, the Foundation became part of the Dana and David Dornsife College of Letters, Arts and Sciences at the University of Southern California in Los Angeles and was renamed as the USC Shoah Foundation Institute for Visual History and Education.</p><br> <p>The goal of the MALACH project was to develop methods for improved access to large multinational spoken archives; the focus was advancing the state of the art of automatic speech recognition and information retrieval. The characteristics of the USC-SFI collection -- unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching and emotional speech -- were considered well-suited for that task. The work centered on five languages: English, Czech, Russian, Polish and Slovak.</p><br> <p>LDC has also released USC-SFI MALACH Interviews and Transcripts Czech (<a href="http://catalog.ldc.upenn.edu/LDC2014S04">LDC2014S04</a>).</p><br> <h3>Data</h3><br> <p>The original MALACH English data set (<a href="../../../LDC2012S05">LDC2012S05</a>) consists of unsegmented audio interviews in mp2 format and speaker-turn, time-marked transcripts in Transcriber (.trs) format presented in a single flat file. In this release, the speech files are segmented and converted to flac format, and the transcripts are updated to an utterance-by-utterance format. Additionally, a lexicon mapping words to phonemes is provided, and the data is divided into development and training sets.</p><br> <p>See the included documentation for more details on these changes, and the documentation and catalog entry for <a href="../../../LDC2012S05">LDC2012S05</a> for further information about the source files.</p><br> <h3>Samples</h3><br> <p>Please view the following samples. Approximately 40 seconds of silence was left at the start of the speech file to preserve the time stamps' accuracy.</p><br> <ul><br> <li><a href="desc/addenda/LDC2019S11.flac">Speech</a></li><br> <li><a href="desc/addenda/LDC2019S11.seg.txt">Segments</a></li><br> <li><a href="desc/addenda/LDC2019S11.txt">Transcript</a></li><br> </ul><br> <h3>Updates</h3><br> <p>None at this time.</p></br> <p>Portions © 2012, 2019 USC Shoah Foundation Institute, © 2012, 2019 Trustees of the University of Pennsylvania</p> <br>The USC-SFI Malach Data is from the archive of the University of Southern California Shoah Foundation Institute for Visual History and Education.</br>

<h3>引言</h3><br><p>USC-SFI MALACH访谈与转录文本 英语版——语音识别专项,语言数据联盟 (LDC) 目录编号LDC2019S11,国际标准书号 (ISBN) 1-58563-889-7,由IBM作为MALACH项目 (Multilingual Access to Large Spoken ArCHives) 的一部分开发。该版本扩充了USC-SFI MALACH访谈与转录文本 英语版(LDC2012S05),针对语音识别工作中的Kaldi工具包 (Kaldi) 修改并更新了原始语料库的一个子集,同时也可轻松移植到其他语音识别系统中使用。本数据集包含来自682名大屠杀幸存者的约168小时访谈音频,配套转录文本、音系词典、适配Kaldi的专用文件及其他文档。</p><br><p>受制作<em>《辛德勒的名单》</em>的经历启发,史蒂文·斯皮尔伯格于1994年创立了大屠杀幸存者视觉历史基金会,旨在收集大屠杀幸存者及其他见证者的视频证词。尽管绝大多数证词提供者为犹太幸存者,但该基金会还采访了同性恋幸存者、耶和华见证人幸存者、解放者及解放见证者、政治犯、救援者与援助提供者、罗姆人与辛提(吉普赛)幸存者、优生政策受害者以及战争罪审判参与者。该基金会的视觉历史档案馆馆藏近5.5万份视频证词,涵盖43种语言、65个国家,是全球同类规模最大的档案馆。2006年,该基金会并入洛杉矶南加州大学达纳与大卫·多恩西夫文理学院,并更名为USC Shoah Foundation视觉历史与教育研究所。</p><br><p>MALACH项目的目标是开发方法,以改进对大型跨国口语档案的获取途径,核心在于推动自动语音识别与信息检索技术的发展。USC-SFI语料库的特性——无约束的自然口语,包含诸多不流畅现象、浓重口音、与年龄相关的协同发音、无提示的说话人与语言切换以及情绪化口语——被认为非常适合该类任务。该项目聚焦五种语言:英语、捷克语、俄语、波兰语与斯洛伐克语。</p><br><p>LDC还发布了USC-SFI MALACH访谈与转录文本 捷克语版(LDC2014S04)。</p><br><h3>数据</h3><br><p>原始MALACH英语数据集(LDC2012S05)包含未分段的mp2格式 (MP2) 音频访谈,以及采用Transcriber格式 (Transcriber) 的、标注了说话人轮次与时间戳的转录文本,存储于单个扁平文件中。在本次发布的版本中,语音文件已被分段并转换为FLAC格式 (FLAC),转录文本也更新为逐语句格式。此外,还提供了单词到音素的映射词典,且数据集被划分为开发集与训练集。</p><br><p>有关这些修改的更多细节,请参阅随附文档;有关源文件的进一步信息,请参阅LDC2012S05的文档与目录条目。</p><br><h3>示例</h3><br><p>请查看以下示例。为保证时间戳的准确性,语音文件开头保留了约40秒的静音。</p><br><ul><br><li><a href="desc/addenda/LDC2019S11.flac">语音</a></li><br><li><a href="desc/addenda/LDC2019S11.seg.txt">分段信息</a></li><br><li><a href="desc/addenda/LDC2019S11.txt">转录文本</a></li><br></ul><br><h3>更新</h3><br><p>暂无更新。</p><br><p>部分内容 © 2012、2019 USC Shoah Foundation视觉历史与教育研究所,© 2012、2019 宾夕法尼亚大学受托人</p><br><br><p>USC-SFI Malach数据集源自南加州大学Shoah Foundation视觉历史与教育研究所的档案馆。</p>

创建时间:
2020-11-30
搜集汇总
数据集介绍
USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition 数据集图片
背景与挑战
背景概述
该数据集是USC-SFI MALACH访谈和转录文本的英语语音识别版,包含约168小时的大屠杀幸存者访谈音频及对应的转录文本,并提供了词典和Kaldi工具包相关文件。它针对口语中的口音、非流利性和语言切换等挑战进行了优化,适用于语音识别系统的研究与开发。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务