nyralabs/disfluency_speech_english
收藏资源简介:
Nyra Disfluency Speech English是一个英语语音数据集,专门用于评估verbatim ASR(逐字自动语音识别)模型。这些模型不仅需要转录说话者的意图词汇,还需包括填充词(如[UH]、[UM])、切断(用*标记,例如th*)、重复以及声音事件(如[laughter]、[breath])。该数据集基于AMAAI Lab的DisfluencySpeech数据集重新格式化,提供了成对的转录:verbatim_transcript(记录说话者实际说的内容,包括所有非流利元素)和intended_transcript(清理后的版本,仅保留说话者的意图意义)。数据集包含4,957个话语,约9.4小时音频,分为训练集(4,458个样本)、验证集(250个样本)和测试集(249个样本)。每个样本包含id、音频、时长、分割集、说话者以及两个转录版本。数据集设计用于Nyra Verbatim Speech Benchmark,以详细评估逐字ASR模型,并分析填充词、声音、切断、重复和意图转录错误。
Nyra Disfluency Speech English is an English speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events. This dataset is based on the AMAAI Lab DisfluencySpeech dataset and reformatted for verbatim-transcription benchmarking with paired verbatim_transcript (what the speaker actually said) and intended_transcript (a cleaned version of what the speaker meant to say). It contains 4,957 utterances and about 9.4 hours of audio, split into train (4,458), validation (250), and test (249) sets. Each sample includes features such as id, audio, duration, split, speaker, and the two transcript versions. It is used by the Nyra Verbatim Speech Benchmark for detailed evaluation of verbatim ASR models, breaking errors down into fillers, sounds, cutoffs, repetitions, and intended-transcript failures.




