RATS Speech Activity Detection
收藏资源简介:
<h3>Introduction</h3><br> <p>RATS Speech Activity Detection was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 3,000 hours of Levantine Arabic, English, Farsi, Pashto, and Urdu conversational telephone speech with automatic and manual annotation of speech segments. The corpus was created to provide training, development and initial test sets for the Speech Activity Detection (SAD) task in the DARPA RATS (Robust Automatic Transcription of Speech) program.</p><br> <p>The goal of the RATS program was to develop human language technology systems capable of performing speech detection, language identification, speaker identification and keyword spotting on the severely degraded audio signals that are typical of various radio communication channels, especially those employing various types of handheld portable transceiver systems. To support that goal, LDC assembled a system for the transmission, reception and digital capture of audio data that allowed a single source audio signal to be distributed and recorded over eight distinct transceiver configurations simultaneously. Those configurations included three frequencies -- high, very high and ultra high -- variously combined with amplitude modulation, frequency hopping spread spectrum, narrow-band frequency modulation, single-side-band or wide-band frequency modulation. Annotations on the clear source audio signal, e.g., time boundaries for the duration of speech activity, were projected onto the corresponding eight channels recorded from the radio receivers.</p><br> <h3>Data</h3><br> <p>The source audio consists of conversational telephone speech recordings collected by LDC: (1) data collected for the RATS program from Levantine Arabic, Farsi, Pashto and Urdu speakers; and (2) material from the Fisher English (<a href="../../../LDC2004S13">LDC2004S13</a>, <a href="../../../LDC2005S13">LDC2005S13</a>), and Fisher Levantine Arabic telephone studies (<a href="../../../LDC2007S02">LDC2007S02</a>), as well as from CALLFRIEND Farsi (<a href="../../../LDC2014S01">LDC2014S01</a>).</p><br> <p>Annotation was performed in three steps. LDC's automatic speech activity detector was run against the audio data to produce a speech segmentation for each file. Manual first pass annotation was then performed as a quick correction of the automatic speech activity detection output. Finally, in a manual second pass annotation step, annotators reviewed first pass output and made adjustments to segments as needed.</p><br> <p>All audio files are presented as single-channel, 16-bit PCM, 16000 samples per second; lossless FLAC compression is used on all files; when uncompressed, the files have typical "MS-WAV" (RIFF) file headers.</p><br> <h3>Samples</h3><br> <p>Please view this <a href="desc/addenda/LDC2015S02.wav">audio sample</a> and <a href="desc/addenda/LDC2015S02.txt">annotation sample</a>.</p><br> <h3>Updates</h3><br> <p>None at this time.</p><br> <h3>Acknowledgment</h3><br> <p>This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. D10PC20016. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.</p></br> Portions © 2015 Trustees of the University of Pennsylvania
<h3>简介</h3><br><p>RATS语音活动检测(Speech Activity Detection, SAD)数据集由语言数据联盟(Linguistic Data Consortium, LDC)开发,包含约3000小时的黎凡特阿拉伯语、英语、波斯语、普什图语与乌尔都语会话电话语音数据,并对语音片段完成了自动与人工标注。本语料库旨在为DARPA RATS(鲁棒语音自动转录,Robust Automatic Transcription of Speech)计划中的语音活动检测任务提供训练集、开发集与初始测试集。</p><br><p>RATS计划的目标是研发可在各类无线电通信信道(尤其是搭载各类手持便携式收发机系统的信道)典型的严重劣化音频信号上,完成语音检测、语言识别、说话人识别与关键词检索任务的人类语言技术系统。为支撑该目标,LDC搭建了一套音频数据传输、接收与数字化采集系统,可将单源音频信号同时分发并录制至8种不同的收发机配置中。这些配置涵盖高、甚高、超高频三类频段,并分别与调幅、跳频扩频、窄带调频、单边带或宽带调频等调制方式组合。将清晰源音频信号的标注信息(如语音活动时段的时间边界),投影至由无线电接收机录制的对应8个通道上。</p><br><h3>数据</h3><br><p>源音频由LDC采集的会话电话语音录音组成:(1) 为RATS计划采集的黎凡特阿拉伯语、波斯语、普什图语与乌尔都语语音数据;(2) 来自Fisher英语语料库(<a href="../../../LDC2004S13">LDC2004S13</a>、<a href="../../../LDC2005S13">LDC2005S13</a>)、Fisher黎凡特阿拉伯语电话语音研究语料库(<a href="../../../LDC2007S02">LDC2007S02</a>)以及CALLFRIEND波斯语语料库(<a href="../../../LDC2014S01">LDC2014S01</a>)的素材。</p><br><p>标注工作分为三个步骤完成。首先,LDC的自动语音活动检测器对音频数据进行处理,为每个文件生成语音分割结果;随后开展第一轮人工标注,对自动语音活动检测的输出结果进行快速修正;最后,在第二轮人工标注阶段,标注人员审阅第一轮输出结果,并根据需要对语音片段进行调整。</p><br><p>所有音频文件均为单通道、16位脉冲编码调制(PCM)格式,采样率为16000样本/秒;所有文件均采用无损FLAC压缩;未压缩时,文件采用标准MS-WAV(RIFF)文件头。</p><br><h3>示例</h3><br><p>请查看该<a href="desc/addenda/LDC2015S02.wav">音频示例</a>与<a href="desc/addenda/LDC2015S02.txt">标注示例</a>。</p><br><h3>更新情况</h3><br><p>暂无更新。</p><br><h3>致谢</h3><br><p>本材料基于美国国防高级研究计划局(Defense Advanced Research Projects Agency, DARPA)在合同号D10PC20016下支持的工作。本内容不一定反映政府的立场或政策,不应被视为获得官方认可。</p><br><p>部分内容 © 2015 宾夕法尼亚大学托管会</p>




