nb-asr-supermorphed-nob
收藏资源简介:
该数据集名为 NbAiLab/nb-asr-supermorphed-nob,旨在将普通挪威博克马尔语(Bokmål)文本直接转换为 AltMorph 括号编码格式。数据集包含训练集(349,692 行,约 901 万源词)、验证集(1,000 行)和测试集(1,000 行),以及一个审计配置(包含被拒绝的行)。每一行保留原始来源数据集、修订版本、行ID、文档ID、来源类型和许可证信息。数据集的构建基于上游训练数据,经过 NFC 和空白字符规范化、按不超过 200 个空白词分割、去重、语言识别过滤(使用挪威语语言识别模型排除非挪威语文本)、语义聚类与采样,并通过规则教师模型和神经网络学生模型生成候选,最终由三位盲审者以多数一致方式接受。自动一致接受率为 94.72%,最终训练集包含 351,556 个自动一致和 369 个经三人审核一致的行。该数据集适用于文本到文本生成任务,特别是形态学编码转换。
The dataset named NbAiLab/nb-asr-supermorphed-nob is designed to convert ordinary Norwegian Bokmål text directly into the AltMorph bracket encoding format. The dataset includes a training set (349,692 rows, approximately 9.01 million source words), a validation set (1,000 rows), a test set (1,000 rows), and an audit configuration (containing rejected rows). Each row retains the original source dataset, revision version, row ID, document ID, source type, and license information. The dataset is constructed based on upstream training data, undergoing NFC and whitespace normalization, splitting by no more than 200 whitespace words, deduplication, language identification filtering (using a Norwegian language identification model to exclude non-Norwegian text), semantic clustering and sampling, and candidate generation through a rule-based teacher model and a neural network student model, finally accepted by three blind reviewers with majority agreement. The automatic consistent acceptance rate is 94.72%, and the final training set contains 351,556 automatically consistent rows and 369 rows agreed upon by three reviewers. This dataset is suitable for text-to-text generation tasks, particularly morphological encoding conversion.




