Nexdata/Burmese_Conversational_Speech_Data_by_Mobile_Phone
收藏资源简介:
该数据集包含120小时的缅甸语对话语音数据,涉及超过130名母语者,性别比例均衡。录音在各种手机上进行,音频格式为16kHz、16bit、未压缩的WAV,所有录音均在安静的室内环境中完成。语音数据包括手动转录的文本内容、每个有效句子的开始和结束时间以及说话者识别。数据集适用于语音识别和声纹识别等应用场景,单词准确率不低于97%。
This dataset contains 120 hours of Burmese conversational speech data, involving over 130 native speakers with a balanced gender ratio. All recordings were conducted on various mobile phones, with the audio format being 16kHz, 16-bit, uncompressed WAV, and all recordings were completed in quiet indoor environments. The speech data includes manually transcribed text content, the start and end timestamps of each valid sentence, and speaker identification information. This dataset is suitable for application scenarios such as speech recognition and speaker verification, with a word accuracy rate of no less than 97%.
数据集卡片 Nexdata/Burmese_Conversational_Speech_Data_by_Mobile_Phone
描述
120小时缅甸语对话语音数据集涉及超过130名母语者,性别比例均衡。参与者从给定列表中选择几个熟悉的话题进行对话,确保对话的流畅性和自然性。录音设备为各种手机,音频格式为16kHz、16bit、未压缩的WAV,所有语音数据在安静的室内环境中录制。所有语音音频均手动转录,包括文本内容、每句有效句子的开始和结束时间以及说话者识别。
规格
格式
16kHz 16bit,未压缩的wav,单声道;
环境
安静的室内环境,无回声;
录音内容
指定数十个话题,说话者在话题下进行对话并录音;
人口统计
共134名说话者,男女各占50%;
标注
转录文本、说话者识别和性别标注;
设备
安卓手机、iPhone;
语言
缅甸语;
应用场景
语音识别;声纹识别;
准确率
单词准确率不低于97%。
许可信息
商业许可




