Speech in Noisy Environments (SPINE2) Part 3 Audio
收藏资源简介:
Introduction Speech in Noisy Environments (SPINE2) Part 3 Audio was used as the evaluation set for the Second Speech in Noisy Environments Evaluation (SPINE2). SPINE2 provided a continuing forum for assessing the state of the art and practice in speech recognition technology for noisy military environments and for exchanging information on innovative speech recognition technology in the context of fully implemented systems that perform realistic tasks. The evaluation provided researchers, potential sponsors, and customers with a quantitative means to appreciate the strengths and weaknesses of the technologies. This work was sponsored in part by National Science Foundation Grant No. IIS-9982201. Data This release contains the Speech in Noisy Environments 2 (SPINE2) Clean and Vocoded Evaluation Audio Corpus created for the Department of Defense (DoD) Digital Voice Processing Consortium (DDVPC) by Arcon Corp. The transcripts for this publication are available as Speech in Noisy Environments (SPINE2) Evaluation Transcripts LDC2001T09. These corpora supported the 2001 Speech in Noisy Environments evaluation. The evaluation data comprises 16 talker pairs (32 speakers total) with four conversations (sessions) per talker pair (64 conversations total). The audio for each session is presented in three forms: Unprocessed: the signal recorded at the participant's microphone Bitstream: the compressed "channel" data produced by the vocoder's analysis stage for transmission from sender to receiver Processed: the signal produced by the vocoder's synthesis stage, given the bitstream data as input. There are a total of 64 clean audio files and 64 vocoded files, one "game" each, for a rough total of seven hours (423 minutes) of audio data, 1.6Gb (including the unprocessed, the processed, and the bitstream files), 23,300 total tokens (930 unique tokens). Samples Please view this unprocessed sample and processed sample. Updates There are no updates at this time.
引言 嘈杂环境语音(Speech in Noisy Environments,SPINE2)第三部分音频被用作第二届嘈杂环境语音评测(Speech in Noisy Environments Evaluation,SPINE2)的评测集。SPINE2搭建了持续交流的专业论坛,用于评估军用嘈杂环境下语音识别技术的前沿水平与实践应用,并针对搭载于可执行真实任务的 fully 部署并投入运行的系统中的创新语音识别技术开展信息交换。本次评测为研究人员、潜在资助方与客户提供了量化评估手段,以研判各类技术的优劣特性。本项目部分由美国国家科学基金会(National Science Foundation)资助,资助编号为IIS-9982201。 数据说明 本次发布包含由Arcon公司为美国国防部(Department of Defense,DoD)数字语音处理联盟(Digital Voice Processing Consortium,DDVPC)制作的《嘈杂环境语音2(SPINE2)纯净与声码化评测音频语料库》。本出版物的转录文本可通过《嘈杂环境语音(SPINE2)评测转录文本》(LDC2001T09)获取。上述语料支撑了2001年嘈杂环境语音评测活动。 本次评测数据包含16组说话人对(总计32名说话人),每组说话人开展4次对话会话,总计64次对话。每次会话的音频以三种形式提供: 未处理版:参与者麦克风录制的原始信号 比特流版:由声码器分析阶段生成、用于从发送端传输至接收端的压缩“信道”数据 处理版:以比特流数据为输入,经声码器合成阶段生成的信号 本次发布共计包含64个纯净音频文件与64个声码化音频文件,每个对应一场“游戏”,总音频时长约7小时(423分钟),总数据量1.6Gb(包含未处理版、处理版与比特流文件),总计23300个Token(930个唯一Token)。 示例 请查看本未处理版音频示例与处理版音频示例。 更新情况 目前暂无更新内容。



