nb-asr-supermorphed-nno
收藏资源简介:
NbAiLab/nb-asr-supermorphed-nno 是一个专为挪威尼诺斯克语(Nynorsk)设计的文本到文本生成数据集,旨在将普通尼诺斯克文本转换为AltMorph括号编码格式。该数据集包含训练集(482,592行,9,457,601个源词)、验证集(1,000行)和测试集(1,000行),此外还有一个审计配置(audit),用于存储被拒绝的样本、候选文本、解码投票及拒绝状态。每个样本都保留了原始来源信息,如源数据集、不可变修订版本、行ID、文档ID、类型和许可证。数据集构建过程严格:采用NFC和空白标准化、去重、语言识别(LID)过滤非挪威语文本、基于E5编码器的语义聚类,以及教师-学生工作流生成候选。规则教师模型使用AltMorph、HumIT标签器、Ordbank和North-T5;学生模型使用T5Gemma 2 1B进行微调。最终通过三重盲审解决分歧,确保高质量。该数据集适用于任何需要从尼诺斯克文本生成AltMorph编码的任务,如文本规范化、形态学转换或机器翻译中的中间表示。
NbAiLab/nb-asr-supermorphed-nno is a text-to-text generation dataset designed specifically for Norwegian Nynorsk. It aims to convert ordinary Nynorsk text into the AltMorph bracket encoding format. The dataset contains a training set (482,592 rows, 9,457,601 source words), a validation set (1,000 rows), and a test set (1,000 rows), plus an audit configuration for storing rejected samples, candidate texts, decoding votes, and rejection status. Each sample retains original source information such as source dataset, immutable revision, row ID, document ID, type, and license. The dataset construction process is rigorous: it employs NFC and whitespace normalization, deduplication, language identification (LID) filtering to remove non-Norwegian text, semantic clustering based on E5 encoder, and a teacher-student workflow to generate candidates. The rule-based teacher model uses AltMorph, HumIT tagger, Ordbank, and North-T5; the student model is fine-tuned T5Gemma 2 1B. Final disagreements are resolved through triple-blind review to ensure high quality. This dataset is suitable for any task requiring generation of AltMorph encoding from Nynorsk text, such as text normalization, morphological transformation, or intermediate representation in machine translation.




