Boxp/ts_asr_test
收藏资源简介:
ts_asr_test是一个目标说话人自动语音识别(Target-Speaker ASR)测试集,仅提供清单(manifest-only),不包含音频文件。数据集包含3928个音频片段,总时长约8.7小时,平均每条8.0秒,采样率为16 kHz单声道。语言包括中文(1580条)和英文(2348条),支持中英混合语音。样本分为正样本(3723条,包含目标说话人、干扰说话人和背景噪声)、负样本(115条,仅噪声/静音)和负样本(90条,仅干扰说话人),用于在给定注册音频(enrollment)的条件下,只转写目标说话人内容的任务。数据集通过确定性的重建脚本从公共源语料库(如LibriSpeech、Common Voice等)生成,支持可复现性。
ts_asr_test is a manifest-only test dataset for Target-Speaker Automatic Speech Recognition (Target-Speaker ASR), which only includes the manifest without any accompanying audio files. It comprises 3,928 audio segments, with a total duration of roughly 8.7 hours and an average length of 8.0 seconds per segment, recorded at 16 kHz mono sampling rate. The dataset covers two languages: Mandarin Chinese (1,580 segments) and English (2,348 segments), and supports mixed Chinese-English speech. Samples are categorized into three groups: positive samples (3,723 segments containing target speakers, interfering speakers, and background noise), and two types of negative samples: 115 segments with only noise or silence, and 90 segments with only interfering speakers. This dataset is developed for the task of transcribing only the target speaker's speech given a reference enrollment audio clip. It is generated from public source corpora including LibriSpeech and Common Voice via deterministic reconstruction scripts, ensuring full reproducibility.




