遇见数据集

Mixer 4 and 5 Speech

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

Description Introduction Mixer 4 and 5 Speech was developed by the Linguistic Data Consortium (LDC) and is comprised of approximately 14,185 hours of audio recordings of conversational telephone speech, interviews, elicitation exercises and transcript readings involving 616 distinct speakers. The material was collected in 2007 as part of the Mixer project and recordings in this corpus were used in the 2008 NIST Speaker Recognition Evaluation (SRE). The data in this release was collected in 2007 by LDC at its Human Subjects Data Collection Laboratories in Philadelphia and by the International Computer Science Institute (ICSI) at the University of California, Berkeley. The Mixer 4 and Mixer 5 collections were conducted simultaneously, as a collaborative, carefully coordinated activity at both recording sites. The telephone protocol connected recruited speakers through a robot operator to carry on casual conversations. In Mixer 4, 400 subjects made ten 10-minute calls; half of those subjects also visited one of the collection sites where they made two telephone calls while also being recorded on a cross-channel platform. In Mixer 5, 300 subjects each completed ten calls and six interview sessions at either LDC or ICSI; those sessions were conducted on a cross channel platform and included a telephone call in one of three vocal-effort conditions - normal, high and low. Mixer participants were nearly all native English speakers, the rest being bilingual English speakers. Researchers interested in applying NIST 2008 SRE benchmark test sets should consult the respective NIST Evaluation Plans for guidelines on allowable training data for those tests. Training, evaluation and supplemental data from 2008 SRE are available in the LDC Catalog: 2008 NIST Speaker Recognition Evaluation Training Set Part 1 (LDC2011S05), 2008 NIST Speaker Recognition Evaluation Training Set Part 2 (LDC2011S07), 2008 NIST Speaker Recognition Evaluation Test Set (LDC2011S08) and 2008 NIST Speaker Recognition Evaluation Supplemental Set (LDC2011S11). Data The Mixer 4 and 5 collection contains 2,568 recordings made via the public telephone network and 2,152 sessions of multiple microphone recordings in office-room settings. The telephone recordings are presented as 8-KHz 2-channel NIST SPHERE files, and the microphone recordings are 16-KHz 1-channel flac/ms-wav files. When the microphone recording flac files are uncompressed, they become ms-wav/RIFF files (flac compression does not presently support SPHERE file format). The telephone audio is presented in SPHERE format because this is consistent with other LDC telephone audio releases and because flac does not support ulaw sample encoding. The open-source SoX utility is able to handle both formats as input. Other utilities are available for flac and SPHERE formats. Metadata about the calls and speakers is also included in this release, along with time-aligned entries for many of the component portions of the recording sessions. Samples Please listen to this telephone sample (SPH) and microphone sample (FLAC). Updates None at this time. Portions © 2007, 2008, 2020 Trustees of the University of Pennsylvania

Mixer 4与Mixer 5语音数据集说明 本数据集由语言数据联盟(Linguistic Data Consortium, LDC)开发,包含约14185小时的语音录音,涵盖对话电话语音、访谈、诱导性练习以及文本朗读内容,涉及616名不同的说话者。 该数据集素材采集于2007年,作为Mixer项目的组成部分,其录音曾被用于2008年美国国家标准与技术研究院(National Institute of Standards and Technology, NIST)说话人识别评测(Speaker Recognition Evaluation, SRE)。 本次发布的数据由LDC于2007年在其位于费城的人类受试者数据采集实验室,以及加州大学伯克利分校的国际计算机科学研究所(International Computer Science Institute, ICSI)共同采集。Mixer 4与Mixer 5的采集工作同步开展,是两个采集站点协同配合、精心统筹的合作项目。 电话通话流程通过机器人话务员连接招募的受试者,开展非正式对话。在Mixer 4项目中,400名受试者完成了10次时长为10分钟的通话;其中半数受试者还到访了其中一个采集站点,在进行两次电话通话的同时,通过跨通道平台完成同步录音。在Mixer 5项目中,300名受试者分别完成了10次通话,以及在LDC或ICSI站点开展的6次访谈环节;所有访谈环节均通过跨通道平台进行,其中包含三种语音力度条件(正常、高声、低声)下的电话通话。 参与Mixer项目的受试者几乎均为英语母语者,剩余少数为英语双语使用者。希望使用2008年NIST SRE基准测试集的研究人员,请参阅相应的NIST评测方案,以了解此类测试允许使用的训练数据相关指南。2008年NIST SRE的训练、评测与补充数据可在LDC目录中获取:2008年NIST说话人识别评测训练集第1部分(LDC2011S05)、2008年NIST说话人识别评测训练集第2部分(LDC2011S07)、2008年NIST说话人识别评测测试集(LDC2011S08)以及2008年NIST说话人识别评测补充集(LDC2011S11)。 数据说明 Mixer 4与Mixer 5数据集包含2568条通过公共电话网络采集的录音,以及2152条办公室环境下多麦克风录制的会话录音。电话录音采用8kHz双声道NIST SPHERE文件格式存储,麦克风录音则采用16kHz单声道FLAC/MS-WAV文件格式。FLAC格式的麦克风录音解压缩后将变为MS-WAV/RIFF文件(当前FLAC压缩格式尚不支持SPHERE文件格式)。电话音频采用SPHERE格式存储,一方面是为了与LDC其他电话语音发布数据集保持格式一致,另一方面是因为FLAC格式不支持μlaw采样编码。开源SoX工具可作为输入处理这两种格式,此外也有其他工具可分别处理FLAC与SPHERE格式。本次发布还包含通话与受试者的元数据,以及录音会话多个组成部分的时间对齐标注条目。 示例 请收听本电话示例(SPH格式)与麦克风示例(FLAC格式)。 更新 暂无更新。 版权说明 部分内容 © 2007、2008、2020 宾夕法尼亚大学董事会

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务