遇见数据集

Fisher English Training Part 2, Transcripts

收藏
Mendeley Data2024-01-31 更新2024-06-27 收录
官方服务:

资源简介:

Introduction Fisher English Training Part 2 Transcripts represents the second half of a collection of conversational telephone speech (CTS) that was created at the LDC during 2003. It consists of time-aligned transcripts for the speech contained in Fisher English Training Part 2, Speech (LDC2005S13). The Fisher telephone conversation collection protocol was created at the LDC to address a critical need of developers trying to build robust automatic speech recognition (ASR) systems. Previous collection protocols, such as CALLFRIEND and Switchboard-II and the resulting corpora, have been adapted for ASR research but were in fact developed for language and speaker identification respectively. Although the CALLHOME protocol and corpora were developed to support ASR technology, they feature small numbers of speakers making telephone calls of relatively long duration with narrow vocabulary across the collection. CALLHOME conversations are challengingly natural and intimate. Under the Fisher protocol, a large number of participants each calls an other participant, whom they typically do not know, for a short short period of time to discuss the assigned topics. This maximizes inter-speaker variation and vocabulary breath while also increasing formality. Previous protocols such as CALLHOME, CALLFRIEND and Switchboard relied upon participant activity to drive the collection. Fisher is unique in being platform driven rather than participant driven. Participants who wish to initiate a call may do so however the collection platform initiates the majority of calls. Participants need only answer their phones at the times they specified when registering for the study. To encourage a broad range of vocabulary, Fisher participants are asked to speak on an assigned topic which is selected at random from a list, which changes every 24 hours and which is assigned to all subjects paired on that day. Some topics are inherited or refined from previous Switchboard studies while others were developed specifically for the Fisher protocol. Data The first half of the collection (Fisher English Training Speech,Part 1) was released by the LDC in 2004 (LDC2004S13 for speech data,LDC2004T19 for transcripts). Taken as a whole, the two parts comprise11,699 recorded telephone conversations. The individual audio files are presented in NIST SPHERE format, and contain two-channel mu-law sample data shorten compression has been applied to all files. Data collection and transcription were sponsored by DARPA and the U.S. Department of Defense, as part of the EARS project for research and development in automatic speech recognition. Samples To see an example of this corpus, please examine this sample. Portions © 2003-2005 Trustees of the University of Pennsylvania

【简介】费舍尔英语训练集第二部分转录文本(Fisher English Training Part 2 Transcripts)是2003年由语言数据联盟(LDC)创建的会话电话语音(CTS)数据集的后半部分。该数据集包含费舍尔英语训练集第二部分语音数据(LDC2005S13)中语音内容的时间对齐转录文本。 费舍尔电话交谈采集协议由LDC开发,旨在满足研发鲁棒性自动语音识别(ASR)系统的开发者的迫切需求。此前的采集协议(如CALLFRIEND、Switchboard-II)及其对应的语料库虽已适配ASR研究,但实际上分别是为语言识别和说话人识别任务开发的。尽管CALLHOME协议及其语料库是为支持ASR技术开发的,但该数据集的特点是说话人数量较少,通话时长相对较长,且整体词汇范围较窄。CALLHOME的交谈内容既自然私密,又对系统颇具挑战。 在费舍尔协议框架下,大量参与者会与一名通常互不相识的其他参与者进行短时通话,讨论指定话题。这一设计最大化了说话人差异与词汇覆盖广度,同时提升了对话的正式程度。此前的协议(如CALLHOME、CALLFRIEND、Switchboard)均依赖参与者主动发起通话来推进采集工作,而费舍尔数据集的独特之处在于,它由采集平台驱动而非参与者驱动。虽允许有意发起通话的参与者自行发起,但绝大多数通话均由采集平台发起。参与者仅需在注册研究时指定的时段接听电话即可。 为覆盖更广泛的词汇范围,费舍尔数据集要求参与者围绕从每日更新的话题列表中随机抽取的指定话题展开交谈,当日配对的所有受试者均使用同一话题。部分话题继承自此前的Switchboard研究并经优化,其余则专为费舍尔协议开发。 【数据】该数据集的前半部分(费舍尔英语训练语音集 第一部分)于2004年由LDC发布(语音数据对应LDC2004S13,转录文本对应LDC2004T19)。两部分合共包含11699段录制的电话交谈内容。单条音频文件采用NIST SPHERE格式存储,为双信道μ律采样数据,且所有文件均应用了Shorten压缩算法。 数据采集与转录工作由美国国防高级研究计划局(DARPA)与美国国防部资助,属于自动语音识别研发相关的EARS项目的一部分。 【示例】若需查看该语料库的示例,请参阅本示例文件。本数据集部分内容 © 2003-2005 宾夕法尼亚大学托管委员会。

创建时间:
2024-01-31
搜集汇总
数据集介绍
Fisher English Training Part 2, Transcripts 数据集图片
背景与挑战
背景概述
该数据集是Fisher English Training Part 2的转录文本部分,包含约975小时英语电话对话的时间对齐转录,旨在支持自动语音识别(ASR)研究。其特点在于采用Fisher协议,通过短时、指定话题的对话收集,以增加词汇多样性和形式性,适用于语音识别技术开发。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务