Aymara-Spanish Parallel Corpus for Neural Machine Translation
收藏官方服务:
资源简介:
Abstract This dataset provides a parallel corpus for the Aymara-Spanish language pair, designed to support research in Neural Machine Translation (NMT) for low-resource languages. The corpus contains aligned sentences. Data Sources & Methodology The data was collected via web scraping from publicly available religious texts (Bible) and digitized educational books. To respect copyright and ensure fair use for research: The sentences have been shuffled randomly, preventing the reconstruction of the original narrative work. The dataset is split into training, validation, and testing subsets. This dataset is strictly for academic and non-commercial research purposes. Contents Aligned text files for Train, Validation, and Test splits.
提供机构:
Zenodo创建时间:
2026-01-09



