遇见数据集

TAN-IBE Aranese corpus

收藏
Zenodo2025-11-24 更新2026-05-26 收录
官方服务:

资源简介:

Aranese corpora created in the project TAN-IBE: Neural machine translation for the Romance languages of the Iberian peninsula (Research project founded by the Spanish Ministry of Science and Innovation PROYECTOS DE GENERACIÓN DE CONOCIMIENTO 2021. MODALIDAD: INVESTIGACIÓN ORIENTADA TIPO B). Corpora from Wikipedia (from the dumps 20250301) wikipedia-oci_aran: monolingual Aranese segments from Wikipedia (1,123 segments) wikipedia-spa-oci_aran.txt: parallel Spanish-Aranese segments found in Wikipedia (17 segments) wikipedia-bactranslated-spa-oci_aran.txt: backtranslated using a Marian neural system. (1,123 segments) Synthetic corpus NTEU-synthetic-spa-oci_aran.txt: Synthetic parallel corpus translating the Spanish part of the NTEU English-Spanish corpus, using Apertium and deleting segments with unknown words. (5,087,903 segments)

提供机构:
Zenodo
创建时间:
2025-11-24
二维码
社区交流群
二维码
科研交流群
商业服务