NbAiLab/nb-asr-mt-filtered
收藏资源简介:
nb-asr-mt-filtered是一个生成的机器翻译数据集,基于NB-ASR translationese管道构建。该数据集包含12种语言对配置,如英语-挪威语(eng-nob)、挪威语-英语(nob-eng)等,每种配置都有训练、验证和测试分割。核心构建规则是反向使用合成数据:当语言X的本地文本被翻译成合成语言Y时,结果行被用作Y到X的训练数据。因此,监督目标始终是本地文本,而源端可能包含翻译痕迹。这个过滤变体应用了保守的多语言双语嵌入完整性过滤器,以去除严重的源-目标不匹配,同时保持广泛覆盖。数据集总行数为10,814,726,用于机器翻译训练和翻译痕迹感知数据构建研究。
nb-asr-mt-filtered is a generated machine-translation manifest dataset built from the NB-ASR translationese pipeline. Each config is a language pair in ISO 639-3 style, such as dan-eng. The core construction rule is reversed use of synthetic data: when native text in language X is translated into synthetic language Y, the resulting row is used as Y-to-X training data. The supervised target is therefore always native text, while the source side may contain translationese artifacts. This filtered variant applies a conservative multilingual bitext embedding sanity filter to remove severe source-target mismatches while preserving broad coverage. The dataset is intended for MT training and research on translationese-aware data construction.




