ISL Meeting Speech Part 1
收藏资源简介:
Introduction ISL Meeting Speech Part 1 was produced by Linguistic Data Consortium (LDC) catalog number LDC2004S05 and ISBN 1-58563-294-5. The ISL Meeting Speech Part 1 is a first subset of the ISL Meeting Corpus (112 meetings). It contains 18 meetings collected at the Interactive Systems Laboratories at Carnegie Mellon University in Pittsburgh, PA during the years 2000-2001. The recorded meetings were either natural meetings where participants needed to meet in the real world, or artificial meetings, which were designed explicitly for the purposes of data collection but still had real topics and tasks. The duration of the meetings in this corpus ranges from eight to 64 minutes and averages at 34 minutes. Word-level orthographic transcriptions are available as ISL Meeting Transcripts Part 1. The transcriptions are available as ISL Meeting Transcripts Part 1. Data The collection includes 105 speech files, for a total of approximately 10 hours of meeting speech. The speech for each meeting consists of wave files for each channel and a wave file containing a mix of all channels. The audio was collected at a 16 kHz sample-rate. Audio files for each meeting are provided as separate time-synchronous recordings for each channel, encoded as 16-bit (little-endian) wave files. During meeting recordings, each speaker wore an individual lapel microphone and was recorded via an Alesis 8-channel mix board and an ECHO Layla 8-channel sound card. This setup was designed to obtain a consumer- or application-style sound quality. All meetings were recorded in the same instrumented meeting area. For an example transcript, please click here. There are a total of 31 unique speakers in the corpus. Meetings involved anywhere from three to nine participants, averaging at five. The corpus contains a significant proportion of non-native English speakers, varying in fluency. Sponsorship The collection and preparation of this corpus was made possible in large part through funding from DARPA, both through the GENOA project and through ROAR. Updates Additional information, updates, bug fixes may be avaibale on the ISL Meeting Room project page. Portions © 2000-2003 Interactive Systems Laboratories, Carnegie Mellon University, Pittsburgh, © 2004 Trustees of the University of Pennsylvania
《ISL会议语音第一部分》(ISL Meeting Speech Part 1)由语言数据联盟(Linguistic Data Consortium, LDC)制作,目录编号为LDC2004S05,ISBN为1-58563-294-5。 ISL会议语音第一部分为ISL会议语料库(共112场会议)的首个子集,收录了2000年至2001年间,于宾夕法尼亚州匹兹堡市卡内基梅隆大学交互系统实验室(Interactive Systems Laboratories)采集的18场会议。本次收录的会议既包含参与者于真实场景中开展的自然会议,也包含为数据采集专门设计但仍采用真实议题与任务的人工会议。该语料库中会议时长介于8至64分钟,平均时长为34分钟。词级正字标注文本可通过《ISL会议标注文本第一部分》(ISL Meeting Transcripts Part 1)获取。 数据详情:本次收录内容包含105个语音文件,总时长约10小时的会议语音。每场会议的语音数据包含各声道独立波形音频文件,以及包含所有声道混音的波形音频文件。音频采集采样率为16 kHz。每场会议的音频文件以各声道独立的时间同步录音形式提供,采用16位(小端序)波形编码格式。会议录制过程中,每位发言者佩戴独立的领夹式麦克风,通过Alesis 8声道混音台与ECHO Layla 8声道声卡进行录音。该录制方案旨在获取消费级或应用级的音频质量。所有会议均在同一配备专业采集设备的标准化会议场地内录制。如需查看标注文本示例,请点击此处。 该语料库共包含31位独立发言者。每场会议的参会人数介于3至9人,平均为5人。语料库中包含相当比例的非英语母语使用者,其英语流利程度各不相同。 资助说明:本语料库的采集与整理工作,主要得益于美国国防高级研究计划局(DARPA)通过GENOA项目与ROAR项目提供的资助。 更新说明:更多相关信息、更新内容与漏洞修复内容,可访问ISL会议项目页面获取。 版权声明:本数据集部分内容 © 2000-2003 匹兹堡卡内基梅隆大学交互系统实验室,© 2004 宾夕法尼亚大学理事会。




