issai/PersonaMix
收藏资源简介:
PersonaMix是一个用于目标说话人自动语音识别和目标存在检测的双语(哈萨克语-英语)基准数据集,针对重叠语音场景发布,与Persona-ASR一起发布。四位说话人(两女两男)每人用哈萨克语和英语朗读11个脚本句子。混合语音包含1-3个干扰说话人,信噪比为-3、0、+3、+6分贝,支持同语言注册(条件A)和跨语言注册(条件B);跨语言条件指在一种语言中注册说话人,并在另一种语言中转录其语音。数据集结构包括混合物目录(按条件分类)、实验JSON文件(包含评估清单,如目标说话人ASR和目标存在检测),以及源录音文件(共88个录音,涵盖所有说话人和语言)。清单字段包括混合物音频、注册音频、转录文本、样本类型、目标说话人ID、语言信息、信噪比和元数据。
PersonaMix is a bilingual (Kazakh-English) benchmark dataset for target speaker automatic speech recognition (ASR) and target presence detection, released for overlapping speech scenarios alongside Persona-ASR. Four speakers (two female, two male) each read 11 scripted sentences in both Kazakh and English. The mixed speech contains 1-3 interfering speakers, with signal-to-noise ratios (SNR) of -3, 0, +3, and +6 dB, supporting both same-language enrollment (Condition A) and cross-language enrollment (Condition B); cross-language enrollment refers to enrolling a speaker in one language and transcribing their speech in the other language. The dataset structure includes mixture directories (classified by condition), experimental JSON files (containing evaluation manifests such as target speaker ASR and target presence detection), and source audio recordings (totaling 88 recordings covering all speakers and languages). The manifest fields include mixture audio, enrollment audio, transcribed text, sample type, target speaker ID, language information, signal-to-noise ratio, and metadata.




