遇见数据集

MODALITY corpus - SPEAKER 10 - SEQUENCE S4

收藏
Mendeley Data2024-01-31 更新2024-06-28 收录
官方服务:

资源简介:

The MODALITY corpus is one of the multimodal database of word recordings in English. It consists of over 30 hours of multimodal recordings. The database contains high-resolution, high-framerate stereoscopic video streams and audio signals obtained from a microphone array and a laptop microphone. The corpus can be employed to develop an AVSR system, as every utterance was labelled. Recordings in noisy conditions can be used to test the robustness of speech recognition systems. The language material was based on a remote control scenario and it includes 231 words -numbers, names of months and days, a set of verbs and nouns related to a computer device control. They were read by speakers as separated words and sequences resulting in a set of 12 recording sessions per speaker. Half of the sessions were recorded in quiet conditions, the other half contained three kinds of intrusive signals (traffic, babble and factory noise). The corpus includes recordings of 42 speakers (33 male, 9 female). The participants include 20 students and staff of Multimedia Systems Department of the Gdańsk University of Technology, 5 students of the Institute of English and American Studies of the University of Gdańsk, and 17 native English speakers. The dataset consist of recordings and visual features for SPEAKER 10: sex: man native speaker: no age: 26 The language material: SEQUENCE S4 All recordings for all speakers are available at http://www.modality-corpus.org/ Sample still from the corpus(SPEAKER 10) Due to the size of the corpus (approx. 2.5 TB of data), every speaker’s recording was placed in a separate zip file of the size approx. 4-7 GB each. The recordings were organized according to the speakers’ language skills. The group A (17 speakers) consists of native-speakers. Non-native speakers recordings (Polish nationals) were placed in the Group B (25 speakers). The audio files use the Waveform Audio File Format (.wav), and contain a single PCM audio stream sampled at 44.1 kSa/s with 16-bit depth. The video files utilize the Matroska Multimedia Container Format (.mkv) in which a video stream in 1080p resolution, captured at 100 fps was placed after being compressed with h.265 codec (using High 4:4:4 profile). The ‘.lab’ files are text files containing the information on word positions in audio files, and follow the HTK label format. Each line of a ‘.lab’ file contains the actual label preceded by start and end times (in 100 ns units) e.g. : 1239620000 1244790000 FIVE which denotes the word “five”, occurring between the 123.962 s and 124.479 s of audio.Word-accurate SNR values calculated for every recording are also included in the ZIP file.

MODALITY语料库是英语单词录音多模态数据库之一,包含超过30小时的多模态录音数据。该数据库收录高分辨率、高帧率的立体视频流,以及由麦克风阵列与笔记本电脑麦克风采集的音频信号。本语料库可用于开发音视频语音识别(Audio-Visual Speech Recognition, AVSR)系统,因所有语音片段均已完成标注;带噪环境下的录音则可用于测试语音识别系统的鲁棒性。该语料库的语言素材基于遥控器控制场景,共涵盖231个单词,包括数字、月份与星期名称,以及一组与计算机设备控制相关的动词和名词。发音者以独立单词或单词序列的形式进行朗读,每位发音者对应12组录音会话;其中6组会话于安静环境下录制,剩余6组则加入了三种干扰信号:交通声、嘈杂人声与工厂环境噪声。该语料库共收录42位发音者的录音,其中男性33名、女性9名。参与录制的发音者包括:格但斯克理工大学多媒体系统系的20名学生与教职工、格但斯克大学英美研究学院的5名学生,以及17名英语母语者。本数据集包含编号为SPEAKER 10的发音者的录音与视觉特征信息:性别为男性,非英语母语者,年龄26岁;语言素材为SEQUENCE S4。所有发音者的录音均可通过http://www.modality-corpus.org/ 获取。语料库样本帧(SPEAKER 10)。由于该语料库总数据量约为2.5TB,每位发音者的录音均打包为独立的ZIP压缩包,单包大小约为4-7GB。录音数据按照发音者的语言能力进行分组:A组(17位发音者)为英语母语者;非母语发音者(波兰籍)的录音被归入B组(25位发音者)。音频文件采用波形音频文件格式(.wav),包含单路PCM音频流,采样率为44.1千次采样每秒(kSa/s),位深度为16比特。视频文件采用Matroska多媒体容器格式(.mkv),其中封装了1080p分辨率、100帧率的视频流,该视频流采用h.265编解码器(High 4:4:4配置档)进行压缩。.lab文件为文本文件,记录了音频文件中单词的位置信息,遵循HTK标注格式。.lab文件的每一行均包含起始时间与结束时间(单位为100纳秒),其后为实际标注内容。例如:1239620000 1244790000 FIVE,表示单词"five"出现在音频的123.962秒至124.479秒区间内。所有录音的逐词精确信噪比(Signal-to-Noise Ratio, SNR)计算结果也一并包含在ZIP压缩包中。

创建时间:
2024-01-31
二维码
社区交流群
二维码
科研交流群
商业服务