AFRIDOC-MT
收藏资源简介:
AFRIDOC-MT是由Masakhane NLP团队创建的一个文档级多语言平行翻译数据集,旨在填补非洲低资源语言在文档级机器翻译领域的空白。该数据集包含605个文档,涵盖健康和信息技术两个领域,每个语言对包含10,000个句子。数据来源于Techpoint Africa和世界卫生组织的英文文章,经过人工翻译和严格的质量控制。AFRIDOC-MT不仅支持英语与非洲语言之间的翻译,还支持非洲语言之间的多向翻译。该数据集的应用领域包括机器翻译模型的训练与评估,特别是在文档级翻译任务中,旨在解决低资源语言在长文档翻译中的一致性和连贯性问题。
AFRIDOC-MT is a document-level multilingual parallel translation dataset developed by the Masakhane NLP team, which aims to fill the research gap in document-level machine translation for low-resource African languages. This dataset contains 605 documents across two domains: healthcare and information technology, with 10,000 sentences per language pair. The source data is extracted from English articles published by Techpoint Africa and the World Health Organization, and has undergone manual translation and strict quality control. AFRIDOC-MT supports not only translation between English and African languages, but also multi-directional translation among various African languages. This dataset can be applied to the training and evaluation of machine translation models, especially for document-level translation tasks, with the core goal of addressing the consistency and coherence issues of low-resource languages in long-document translation scenarios.




