OLDI Seed Corpus French Partition
收藏资源简介:
OLDI Seed Corpus French Partition是一个法语分区,由Inria研究机构创建,旨在解决低资源语言翻译训练数据不足的问题。该数据集包含大约6000个英文句子,这些句子从维基百科的核心文章中抽取,涵盖了广泛的主题。为了创建这个数据集,使用了多种机器翻译系统和定制的后编辑界面,由母语为法语的专业人员进行后编辑。这个法语语料库不仅是翻译的终点,而且作为关键的中转资源,旨在促进法国低资源区域语言的平行语料库的收集。该数据集以CC BY-SA 4.0许可证公开可用。
OLDI Seed Corpus French Partition is a French partition created by the Inria research institute, aiming to address the shortage of training data for low-resource language translation. This dataset contains approximately 6,000 English sentences extracted from core Wikipedia articles, covering a wide range of topics. To develop this dataset, multiple machine translation systems and a custom post-editing interface were used, with post-editing performed by professional native French speakers. This French corpus not only serves as a target translation resource but also acts as a key intermediate resource intended to facilitate the collection of parallel corpora for low-resource regional languages in France. This dataset is publicly available under the CC BY-SA 4.0 license.




