TAN-IBE Aragonese corpus
收藏资源简介:
Aragonese corpora created in the project TAN-IBE: Neural machine translation for the Romance languages of the Iberian peninsula (Research project founded by the Spanish Ministry of Science and Innovation PROYECTOS DE GENERACIÓN DE CONOCIMIENTO 2021. MODALIDAD: INVESTIGACIÓN ORIENTADA TIPO B). All the corpora are in the orthographic norm of the Academia Aragonesa de la Lengua: Corpora from Wikipedia (from the dumps 20250301) wikipedia-AAL-arg.txt: monolingual Aragonese (349,069 segments) wikipedia-AAL-spa-arg.txt: parallel Spanish-Aragonese segments found in Wikipedia (6,051 segments) wikipedia-AAL-bactranslated-nounk-spa-arg.txt: backtranslated using wikipedia-AAL-arg.txt with Apertium deleting segments with unknown words. (52,810 segments) Synthetic corpus NTEU-AAL-synthetic-AAL-spa-arg.txt: Synthetic parallel coprus translating the Spanish part of the NTEU English-Spanish corpus, using Apertium and deleting segments with unknown words. (5,019,629 segments)



