nob-nno-eng-translation-pairs
收藏资源简介:
该数据集专为挪威语-英语机器翻译任务设计,旨在用于微调大型语言模型(LLMs)的多句子翻译能力。数据集主要来源于CCAligned,通过提取对齐文档中的连续文本段并进行严格筛选,包括表面级启发式过滤、基于jinaai/jina-embeddings-v3的语义相似性检查,以及使用meta-llama/Llama-3.3-70B-Instruct作为评判模型的匹配验证。此外,数据集还包含了来自NorSumm的Bokmål-Nynorsk人工翻译和Tatoeba开发集的翻译内容。数据集适用于机器翻译任务,特别是挪威语与英语之间的翻译。
This dataset is specifically designed for Norwegian-English machine translation tasks, aiming to fine-tune the multi-sentence translation capabilities of Large Language Models (LLMs). It is primarily sourced from CCAligned, where continuous text segments are extracted from aligned documents and subjected to rigorous filtering, including surface-level heuristic filtering, semantic similarity checks based on jinaai/jina-embeddings-v3, and matching validation using meta-llama/Llama-3.3-70B-Instruct as the judging model. Additionally, the dataset contains manually translated Bokmål-Nynorsk content from NorSumm and translation samples from the Tatoeba development set. This dataset is suitable for machine translation tasks, with a particular focus on Norwegian-English translation.




