Mixer 7 Spanish Speech
收藏资源简介:
Mixer 7 Spanish Speech (LDC2023S04) was developed by the Linguistic Data Consortium (LDC) and contains 9,600 hours of audio recordings of interviews, transcript readings and conversational telephone speech involving 191 distinct native Spanish speakers. This material was collected by LDC in 2011 and 2012 as part of the Mixer project. The recordings in this corpus were used in the 2012 NIST Speaker Recognition Evaluation test set. The speech data in this release was collected by LDC at its Human Subjects Data Collection Laboratories in Philadelphia. The telephone collection protocol was similar to other LDC Mixer collections: recruited speakers were connected through a robot operator to carry on casual conversations lasting up to 10 minutes, usually about a daily topic announced by the robot operator at the start of the call. The raw digital audio content for each call side was captured as a separate channel, and each full conversation is presented as a 2-channel interleaved audio file, with 8000 samples/second and u-law sample encoding. Each speaker was asked to complete 15 calls. The multi-microphone portion of the collection utilized 14 distinct microphones installed identically in two multi-channel audio recording rooms at LDC. Each session was guided by collection staff using prompting and recording software to conduct the following activities: (1) repeat questions (less than one minute); (2) informal conversation (typically 15 minutes); (3) transcript reading (15 minutes); and (4) up to three telephone calls under varying conditions (10 minutes). The 14 channels were recorded synchronously into separate single-channel files, using 16-bit PCM sample encoding at 16000 samples/second. Certain demographic information about the speakers was collected, including date of birth, level of education, native language, other language capability, place of birth, place of residence and occupation. The collection contains 2,583 recordings made via the public telephone network and 678 sessions of multiple microphone recordings in office room settings. The telephone recordings are presented as 8-KHz 2-channel NIST SPHERE files, and the microphone recordings are 16-KHz 1-channel flac/ms-wav files. When the flac files are uncompressed, they become ms-wav/RIFF files (flac compression does not presently support SPHERE file format). The telephone audio is presented in SPHERE format because (a) this is consistent with other telephone audio releases from LDC, and (b) flac does not support ulaw sample encoding. The current release of the open-source SoX utility is able to handle both formats as input. Other utilities are available for both flac and SPHERE formats.
Mixer 7 Spanish Speech (LDC2023S04) 由语言数据联盟(Linguistic Data Consortium, LDC)开发,包含总计9600小时的音频录音,涵盖访谈、文本朗读与会话电话语音三类内容,涉及191名不同的西班牙语母语使用者。该数据集于2011年和2012年作为Mixer项目的一部分由LDC完成采集,其中的录音曾被用于2012年NIST(美国国家标准与技术研究院)说话人识别评测测试集。本次发布的语音数据由LDC在其费城的人类受试者数据采集实验室采集。其电话语音采集流程与LDC此前的其他Mixer系列采集项目一致:招募的说话者通过机器人接线员接入通话,开展最长10分钟的日常闲聊,通话主题通常由接线员在通话开始时公布。每一方通话的原始数字音频将作为独立声道采集,完整会话以2声道交错音频文件形式呈现,采样率为8000样本/秒,采用u-law(μ律)采样编码。每位说话者需完成15通通话。本次采集的多麦克风部分使用了14台安装参数一致的麦克风,分别部署在LDC的两间多声道录音室内。每场会话由采集人员通过提示与录音软件引导完成以下四项活动:(1) 重复提问环节(时长不超过1分钟);(2) 非正式会话(通常持续15分钟);(3) 文本朗读(时长15分钟);(4) 三种不同条件下的电话通话(单次最长10分钟)。14个声道将以同步方式录制为独立单声道文件,采用16位脉冲编码调制(PCM)采样编码,采样率为16000样本/秒。采集过程中还收集了说话者的多项人口统计信息,包括出生日期、教育水平、母语、其他语言能力、出生与居住地点以及职业。本次采集包含2583通通过公共电话网络录制的音频,以及678场办公室环境下的多麦克风录音。其中电话录音以8kHz 2声道NIST SPHERE格式呈现,麦克风录音则为16kHz 1声道FLAC/ms-wav格式。FLAC文件解压后将转换为ms-wav/RIFF格式(目前FLAC压缩尚不支持SPHERE文件格式)。电话音频采用SPHERE格式的原因在于:(a) 该格式与LDC此前发布的其他电话音频数据集格式保持一致;(b) FLAC不支持u-law(μ律)采样编码。当前开源的SoX工具可作为输入处理这两种格式,同时也有其他工具可用于处理FLAC与SPHERE格式。



