nyralabs/disfluency_speech_german
收藏资源简介:
Nyra Disfluency Speech German 是一个德语语音数据集,专门用于评估逐字自动语音识别(verbatim ASR)模型,要求模型不仅能转录说话者的意图单词,还能准确转录填充词、中断、重复和声音事件等不流利内容。该数据集由Nyra研究人员Berns和Laurin内部录制,旨在模拟自然的德语不流利语音,类似于英语的AMAAI Lab DisfluencySpeech数据集。数据集包含202个话语,总计约0.95小时音频,分为测试集。每个样本包括id、音频、时长、分割、说话者、语言、逐字转录(verbatim_transcript)和意图转录(intended_transcript)。逐字转录遵循特定约定,如使用*标记中断(例如w*)、括号标签标记填充词(如[UH]和[UM])和声音事件(如[lipsmack]、[throatclearing]),而意图转录则清理了这些不流利内容,保留说话者的意图。数据集设计用于与Nyra Verbatim Speech Benchmark配合,以详细评估逐字ASR模型的性能,包括计算vWER和iWER等指标,以及分析填充词、声音、中断和重复的错误。统计显示,所有话语都至少包含一个标记事件或中断,填充词和中断频繁出现,使其成为评估不流利语音转录的核心资源。
`nyrahealth/disfluency_speech_german` is a German speech dataset for evaluating verbatim ASR: models that should transcribe not only the intended words, but also fillers, cutoffs, repetitions, and sound events. This dataset was recorded in-house by two Nyra researchers, Berns and Laurin, with the goal of producing natural disfluent German speech similar in spirit to the English AMAAI Lab DisfluencySpeech dataset. It is formatted for verbatim-transcription benchmarking with paired verbatim_transcript and intended_transcript. The dataset contains 202 utterances and about 0.95 hours of audio, with features including id, audio, duration_in_s, split, speaker, language, verbatim_transcript, and intended_transcript. It is used by the Nyra Verbatim Speech Benchmark for detailed evaluation of verbatim ASR, breaking errors down into fillers, sounds, cutoffs, repetitions, and intended-transcript failures. Statistics show every utterance contains at least one annotated event or cutoff, making it dense in disfluencies for core evaluation.




