NbAiLab/nb-asr-mt-gold
收藏资源简介:
nb-asr-mt-gold是一个基于NB-ASR翻译流程生成的机器翻译数据集,包含12种语言对配置(如英语-挪威语、挪威语-英语、瑞典语-英语等)。数据集的构建规则是反向使用合成数据:当语言X的母语文本被翻译成合成语言Y时,结果行被用作Y到X的训练数据,因此监督目标始终是母语文本文本,而源端可能包含翻译痕迹(即非母语表达)。该数据集是nb-asr-mt的变体,通过多语言嵌入模型进行严格筛选,确保母语原文和回译文本在语义上高度相似,旨在提供高置信度的训练数据,覆盖有用内容而非完全相同的字符串。数据集总行数为3,397,227,包含训练、验证和测试分割,用于机器翻译训练和翻译痕迹感知的数据构建研究。
nb-asr-mt-gold is a generated machine-translation manifest dataset built from the NB-ASR translationese pipeline. It includes 12 language pair configurations (e.g., eng-nob, nob-eng, swe-eng). The core construction rule involves reversed use of synthetic data: when native text in language X is translated into synthetic language Y, the resulting row is used as Y-to-X training data, so the supervised target is always native text, while the source side may contain translationese artifacts. This gold variant is filtered using a multilingual embedding model to ensure high similarity between the native original and backtranslation, targeting useful high-confidence coverage rather than exact string equality. The dataset contains 3,397,227 rows with train, validation, and test splits, intended for MT training and research on translationese-aware data construction.




