Helsinki-NLP/tatoeba_mt
收藏资源简介:
Tatoeba翻译挑战是一个多语言机器翻译基准数据集,数据来源于Tatoeba.org用户贡献的翻译,并通过OPUS平台整理成平行语料库。该数据集包含数百种语言对的测试和开发数据,并持续更新。数据集的结构为TAB分隔的文件,包含源语言和目标语言的ISO-639-3代码、源语言文本和目标语言文本。数据集的目标是提供高语言覆盖率的测试集,适用于低资源语言和多语言机器翻译任务。
Tatoeba Translation Challenge is a multilingual machine translation benchmark dataset. The data is sourced from user-contributed translations on Tatoeba.org, and compiled into parallel corpora via the OPUS platform. This dataset includes test and development data for hundreds of language pairs, and is continuously updated. The dataset is structured as TAB-separated files, containing ISO-639-3 codes for both source and target languages, as well as the source language text and target language text. The goal of this dataset is to provide a test set with high linguistic coverage, suitable for low-resource language and multilingual machine translation tasks.
数据集概述
数据集名称
- 名称: The Tatoeba Translation Challenge
- 别名: Tatoeba MT Challenge
数据集内容
- 类型: 机器翻译基准数据集
- 来源: 用户贡献的翻译,由Tatoeba.org收集并由OPUS提供为平行语料库
- 覆盖语言: 数百种语言和语言对,包括但不限于Afrikaans, Arabic, Azerbaijani等
数据集结构
- 数据实例: 翻译单元,以TAB分隔的文件形式,包含源语言和目标语言ISO-639-3代码、源语言文本和目标语言文本
- 数据分割: 测试和开发数据集,测试集最多包含10,000个实例
数据集创建
- 数据收集: 从Tatoeba.org用户贡献的翻译中收集
- 数据准备: 持续更新,数据准备过程公开并发布在GitHub上
- 注释过程: 由志愿者进行翻译,注释者包括各种语言技能的贡献者
使用许可
- 许可类型: CC-BY 2.0
数据集用途
- 任务类型: 条件文本生成
- 任务ID: 机器翻译
数据集维护
- 维护者: 赫尔辛基大学语言技术研究组
- 维护平台: OPUS生态系统
引用信息
- 引用文献: The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT
- 引用格式:
@inproceedings{tiedemann-2020-tatoeba, title = "The Tatoeba Translation Challenge {--} Realistic Data Sets for Low Resource and Multilingual {MT}", author = {Tiedemann, J{"o}rg}, booktitle = "Proceedings of the Fifth Conference on Machine Translation", month = nov, year = "2020", address = "Online", publisher = "Association for Computational Linguistics", url = "https://aclanthology.org/2020.wmt-1.139", pages = "1174--1182", }




