nb-asr-numerics-balanced-disfluent2
收藏资源简介:
该数据集名为Norwegian ASR Numeric Spelling Balanced Disfluent 2,是一个专为鲁棒声学和语音模型训练而设计的挪威语数据集。它包含1,500,000条经过平衡的句子,这些句子融合了数字和自然的不流畅现象,如填充词(例如ee…、ehm…、mm…、mhm)以及单词或短语的重复。数据集的目的是提供干净的、已验证的输入文本及其对应的、包含自然不流畅性的逐字输出文本。数据集包含三个字段:id(以dis-为前缀的唯一句子标识符)、text(经过数字移位处理的挪威博克马尔语句子,其中数字4与5、6与7、8与9进行了配对交换,而0、1、2、3保持不变,以保持时间、日期等结构的真实性和防止模型记忆)和text_verbatim(经过验证的、包含自然不流畅性的句子版本)。数据集的构建经过了严格的流程:首先从源数据集中选择和平衡包含稀有类别和高数字密度的句子;然后添加ID前缀;接着对文本中的数字进行配对交换移位;最后使用特定的生成模型和提示词生成不流畅文本,并通过一个严格的验证门控确保移除填充词和重复后能精确重构原始单词序列,未通过验证的行被丢弃。最终,在1,500,000条基行中,成功生成了1,427,532行数据,成功率为95.17%。该数据集适用于自动语音识别(ASR)、文本到语音(TTS)以及需要处理数字和口语不流畅现象的语音相关任务。
The dataset is named Norwegian ASR Numeric Spelling Balanced Disfluent 2 and is a Norwegian language dataset designed for robust acoustic and speech model training. It contains 1,500,000 balanced sentences that incorporate numbers and natural disfluencies, such as fillers (e.g., ee…, ehm…, mm…, mhm) and repetitions of words or phrases. The purpose of the dataset is to provide clean, verified input text and its corresponding verbatim output text containing natural disfluencies. The dataset includes three fields: id (a unique sentence identifier prefixed with dis-), text (Norwegian Bokmål sentences with numeric shift processing, where digits 4 and 5, 6 and 7, 8 and 9 are swapped in pairs, while 0, 1, 2, 3 remain unchanged to preserve the authenticity of structures like time and dates and prevent model memorization), and text_verbatim (the verified version of the sentence with natural disfluencies). The construction process involves a rigorous workflow: first, selecting and balancing sentences from the source dataset that include rare categories and high digit density; then adding ID prefixes; next, applying paired digit swapping shifts to the text; finally, using a specific generation model and prompts to generate disfluent text, with a strict validation gate ensuring exact reconstruction of the original word sequence after removing fillers and repetitions, and discarding rows that fail validation. Ultimately, out of 1,500,000 base rows, 1,427,532 rows were successfully generated, with a success rate of 95.17%. This dataset is suitable for automatic speech recognition (ASR), text-to-speech (TTS), and other speech-related tasks that require handling numbers and spoken disfluencies.




