TAN-IBE Asturian corpus
收藏资源简介:
Asturian corpora created in the project TAN-IBE: Neural machine translation for the Romance languages of the Iberian peninsula (Research project founded by the Spanish Ministry of Science and Innovation PROYECTOS DE GENERACIÓN DE CONOCIMIENTO 2021. MODALIDAD: INVESTIGACIÓN ORIENTADA TIPO B). Corpora from Wikipedia (from the dumps 20250301) wikipedia-ast.txt: monolingual Asturian (2.333.988 segments) wikipedia-spa-ast.txt: parallel Spanish-Asturian segments found in Wikipedia (1,539,239 segments) wikipedia-bactranslated-spa-ast.txt: backtranslated using a Marian neural system. (2,333,847 segments) Synthetic corpus NTEU-synthetic-spa-ast.txt: Synthetic parallel corpus translating the Spanish part of the NTEU English-Spanish corpus, using Apertium and deleting segments with unknown words. (7,220,441segments)



