somosnlp-hackathon-2022/Axolotl-Spanish-Nahuatl
收藏资源简介:
Axolotl Spanish-Nahuatl平行语料库是一个数字语料库,汇集了多种来源的西班牙语和纳瓦特尔语平行文本。平行语料库是一种包含源语言文本及其在一种或多种目标语言中的对应翻译的语料库。该语料库由墨西哥国立自治大学(UNAM)的专家团队创建,旨在支持西班牙语和纳瓦特尔语之间的机器翻译任务。数据集来源于Axolotl和Bible UEDIN Nahuatl Spanish语料库,经过清理和去重后,最终包含20,028个样本。数据集的应用包括使用T5模型进行西班牙语到纳瓦特尔语的翻译任务。
The Axolotl Spanish-Nahuatl Parallel Corpus is a digital corpus that aggregates parallel texts in Spanish and Nahuatl from multiple sources. A parallel corpus is a corpus containing source language texts and their corresponding translations in one or more target languages. This corpus was developed by a team of experts from the National Autonomous University of Mexico (UNAM) to support machine translation tasks between Spanish and Nahuatl. The corpus is derived from the Axolotl and Bible UEDIN Nahuatl Spanish corpora, and after cleaning and deduplication, it ultimately contains 20,028 samples. Applications of this corpus include Spanish-to-Nahuatl machine translation tasks using the T5 model.
Axolotl-Spanish-Nahuatl 数据集概述
数据集描述
- 名称: Axolotl Spanish-Nahuatl 平行语料库
- 类型: 平行语料库,包含西班牙语和纳瓦特尔语的平行文本
- 语言:
- 源语言: 西班牙语 (es)
- 目标语言: 纳瓦特尔语
- 许可证: MPL-2.0
- 多语言性: 翻译
- 任务类别:
- 文本到文本生成
- 翻译
- 数据集来源: 原始数据
- 数据集大小: 未知
- 数据集收集:
- 来源1: Axolotl,由UNAM的专家团队收集
- 来源2: Bible UEDIN Nahuatl Spanish,由Christos Christodoulopoulos和Mark Steedman从Bible Gateway网站爬取
- 数据集样本数:
- Axolotl: 12,207样本
- Bible UEDIN: 7,821样本
- 总计: 20,028条语音
团队成员
- Emilio Morales (milmor)
- Rodrigo Martínez Arzate (rockdrigoma)
- Luis Armando Mercado (luisarmando)
- Jacobo del Valle (jjdv)
应用
- 模型: 西班牙语到纳瓦特尔语翻译任务,使用T5模型 (t5-small-spanish-nahuatl)
- 演示: 西班牙语到纳瓦特尔语翻译 (Spanish-nahuatl)




