NbAiLab/nb-asr-numerics-balanced-disfluent
收藏资源简介:
该数据集提供了一个类别平衡的挪威博克马尔语句子合成语料库,包含数字表达,专门为具有自然不流利现象(如填充词和犹豫)的鲁棒自动语音识别(ASR)训练而准备。它从59个数字类别中每个类别抽取20,000个示例(总计超过110万行),使用随机种子43生成数字表达。更新内容包括添加了三个关键字段:tts_text(将数字表达式和数学缩写完全写为自然挪威语单词)、tts_text_disfluent(包含自然填充词和单词重复的轻度不流利版本句子)和text_vebatim(不流利的口语文本,但所有数字表达式以原始文本中的数字/符号形式写出)。不流利句子的生成遵循严格的结构约束,以确保语义等价性,包括仅允许填充词和最多3个单词的重复,保留原始单词顺序和内容,并通过正则表达式过滤器和单词对齐检查进行验证。数据集的模式细节包括唯一标识符、文本字段、实体列表和约束等。
This dataset provides a class-balanced synthetic corpus of Norwegian Bokmål sentences containing numeric expressions, prepared specifically for robust ASR training with natural disfluencies (fillers and hesitations). It draws 20,000 examples for each of the 59 numeric categories (totaling over 1.1 million rows) with randomized seed 43 numeric expressions. The dataset includes updates such as the addition of key fields: tts_text (numeric expressions and mathematical abbreviations fully written out as natural Norwegian words), tts_text_disfluent (a natural, lightly disfluent version with fillers and word repetitions), and text_vebatim (the disfluent spoken text with numeric expressions written as digits/symbols). Disfluency generation is constrained to allow only filler words and repetitions up to 3 words, preserving word order and content, with strict gatekeeper validation via regex filters and word alignment checks. Schema details cover unique identifiers, text fields, entity lists, and constraints.




