Neurora/versta-tonality-en-nl
收藏资源简介:
Versta Tonality EN-NL是一个用于微调的合成翻译数据集。它从OpenCorpus(包括WikiMatrix、TED2020、Europarl、CCAligned、OpenSubtitles、Tatoeba、NLLB、GlobalVoices、kNews-Commentary)中,通过本地托管的教师模型蒸馏而来。该数据集专为Android设备的边缘部署设计,支持通过训练后的学生模型实现离线、低延迟的机器翻译。整个数据集使用可再生能源生成。数据集旨在弥合高精度神经机器翻译与边缘推理速度限制之间的差距,通过将语言知识从教师模型蒸馏到紧凑的LLM架构中,使微调后的学生模型在保持实时移动应用所需亚秒级延迟的同时,实现与更大模型相媲美的翻译质量。数据集由Ricardo Snoek-Valkenburg策划,语言为英语和荷兰语,采用Creative Commons Attribution Share Alike 4.0 (CC-BY-SA-4.0)许可证。
Versta Tonality EN-NL is a synthetic translation dataset for fine-tuning. A multilingual translation dataset distilled from OpenCorpus (WikiMatrix, TED2020, Europarl, CCAligned, OpenSubtitles, Tatoeba, NLLB, GlobalVoices, kNews-Commentary) using a locally-hosted teacher model. Designed for edge deployment on Android devices, this dataset enables offline, low-latency machine translation via trained students. Entirely generated using renewable energy. This dataset was created to bridge the gap between high-accuracy neural machine translation and the speed constraints of edge inference. By distilling linguistic knowledge from a locally-hosted teacher model into the compact LLM architecture, it enables a fine-tuned student model that achieves translation quality rivaling larger models while maintaining the sub-second latency required for real-time mobile applications. Curated by Ricardo Snoek-Valkenburg, language(s) include English and Dutch, with a Creative Commons Attribution Share Alike 4.0 (CC-BY-SA-4.0) license.




