Feriji Dataset
收藏资源简介:
Feriji Dataset是一个专为机器翻译任务设计的法语-扎尔马语平行语料库,由阿什西大学和罗切斯特理工学院的研究团队创建。该数据集包含61,085条扎尔马语句子和42,789条法语句子,所有句子均进行了精确的对齐处理。数据集的创建过程涉及广泛的数据收集、对齐和清洗,确保了数据的质量和适用性。Feriji Dataset不仅填补了扎尔马语在机器翻译领域的资源空白,还促进了这一非洲本土语言在研究领域的应用,旨在解决语言资源不足和促进文化传承的问题。
Feriji Dataset is a French-Zarma parallel corpus specifically designed for machine translation tasks, created by research teams from Ashi University and Rochester Institute of Technology. This dataset contains 61,085 Zarma sentences and 42,789 French sentences, with all sentence pairs being precisely aligned. The dataset creation process involves extensive data collection, alignment and cleaning, ensuring data quality and applicability. The Feriji Dataset not only fills the resource gap of Zarma in the field of machine translation, but also promotes the application of this African indigenous language in research, aiming to address the issue of insufficient language resources and facilitate cultural heritage preservation.

- 1Feriji: A French-Zarma Parallel Corpus, Glossary & Translator阿什西大学 罗切斯特理工学院 · 2024年



