MODALITY corpus - SPEAKER 22 - COMMANDS C1
收藏资源简介:
The MODALITY corpus is one of the multimodal database of word recordings in English. It consists of over 30 hours of multimodal recordings. The database contains high-resolution, high-framerate stereoscopic video streams and audio signals obtained from a microphone array and a laptop microphone. The corpus can be employed to develop an AVSR system, as every utterance was labelled. Recordings in noisy conditions can be used to test the robustness of speech recognition systems. The language material was based on a remote control scenario and it includes 231 words -numbers, names of months and days, a set of verbs and nouns related to a computer device control. They were read by speakers as separated words and sequences resulting in a set of 12 recording sessions per speaker. Half of the sessions were recorded in quiet conditions, the other half contained three kinds of intrusive signals (traffic, babble and factory noise). The corpus includes recordings of 42 speakers (33 male, 9 female). The participants include 20 students and staff of Multimedia Systems Department of the Gdańsk University of Technology, 5 students of the Institute of English and American Studies of the University of Gdańsk, and 17 native English speakers. The dataset consist of recordings and visual features for SPEAKER 22: sex: man native speaker: yes age: 51 The test material: COMMANDS C1 All recordings for all speakers are available at http://www.modality-corpus.org/ Sample still from the corpus(SPEAKER 22) Due to the size of the corpus (approx. 2.5 TB of data), every speaker’s recording was placed in a separate zip file of the size approx. 4-7 GB each. The recordings were organized according to the speakers’ language skills. The group A (17 speakers) consists of native-speakers. Non-native speakers recordings (Polish nationals) were placed in the Group B (25 speakers). The audio files use the Waveform Audio File Format (.wav), and contain a single PCM audio stream sampled at 44.1 kSa/s with 16-bit depth. The video files utilize the Matroska Multimedia Container Format (.mkv) in which a video stream in 1080p resolution, captured at 100 fps was placed after being compressed with h.264 codec (using High 4:4:4 profile). The ‘.lab’ files are text files containing the information on word positions in audio files, and follow the HTK label format. Each line of a ‘.lab’ file contains the actual label preceded by start and end times (in 100 ns units) e.g. : 1239620000 1244790000 FIVE which denotes the word “five”, occurring between the 123.962 s and 124.479 s of audio.Word-accurate SNR values calculated for every recording are also included in the ZIP file.
MODALITY语料库(MODALITY corpus)是一款面向英语词汇录制的多模态数据库。该语料库包含超过30小时的多模态录制数据。数据库内包含高分辨率、高帧率的立体视频流,以及由麦克风阵列和笔记本电脑麦克风采集的音频信号。由于每一段语音都已标注,该语料库可用于开发音视频语音识别(Audio-Visual Speech Recognition, AVSR)系统。带噪环境下的录制数据可用于测试语音识别系统的鲁棒性。 该语料的语言素材基于遥控器操控场景,涵盖231个词汇——包括数字、月份与星期名称,以及一组与计算机设备操控相关的动词和名词。录制时,参与者分别朗读单个词汇与词汇序列,每位参与者共完成12组录制会话。其中6组会话在安静环境下录制,剩余6组则添加了三种干扰信号:交通声、人声嘈杂声与工厂环境噪声。 该语料库共收录42位参与者的录制数据,其中男性33名,女性9名。参与者构成如下:20名来自格但斯克工业大学(Gdańsk University of Technology)多媒体系统系的学生与教职工,5名来自格但斯克大学(University of Gdańsk)英美研究学院的学生,以及17名以英语为母语的参与者。 本数据集包含编号为SPEAKER 22的参与者的录制数据与视觉特征:性别为男性,母语为英语,年龄51岁。测试素材:命令集C1。 所有参与者的全部录制数据均可通过以下网址获取:http://www.modality-corpus.org/。语料库(SPEAKER 22)样本帧。 由于该语料库总数据量约为2.5TB,每位参与者的录制数据均打包为独立的ZIP压缩包,单包大小约为4至7GB。录制数据按照参与者的语言能力分为两组:A组(17名参与者)为以英语为母语的参与者,B组(25名参与者)为非母语参与者(均为波兰籍)。 音频文件采用波形音频文件格式(Waveform Audio File Format, .wav),包含单路脉冲编码调制(Pulse Code Modulation, PCM)音频流,采样率为44.1千采样每秒(kSa/s),位深度为16比特。视频文件采用Matroska多媒体容器格式(Matroska Multimedia Container Format, .mkv),其中的视频流分辨率为1080p,采集帧率为100帧每秒(fps),并采用h.264编解码器(High 4:4:4配置档)进行压缩。 .lab文件为文本文件,记录了音频文件中词汇的位置信息,遵循隐马尔可夫模型工具包(Hidden Markov Model Toolkit, HTK)标注格式。.lab文件的每一行均包含起始时间与结束时间(单位为100纳秒)以及对应的实际标注,示例如下:1239620000 1244790000 FIVE,该标注表示单词"five"出现在音频的123.962秒至124.479秒区间内。此外,ZIP压缩包中还包含为每一段录制数据计算得到的词级精确信噪比(Signal-to-Noise Ratio, SNR)数值。



