Spanish–Aymara Biblical Parallel Corpus v2.0: 36,059 High-Fidelity Sentence Pairs
收藏资源简介:
Description This dataset is an enhanced Spanish–Aymara parallel corpus specifically designed for Neural Machine Translation (NMT) and linguistic research on polysynthetic languages. It builds upon and extends the previous dataset by Nina et al. (2025), introducing significant improvements in data quality, segmentation, and experimental integrity. Dataset Composition The corpus consists of 36,059 high-quality sentence pairs sourced from public biblical texts. The biblical domain was selected for its high morphological fidelity, as these translations are typically produced by native speakers, ensuring a reliable representation of Aymara's complex grammatical structures. Key Enhancements over Nina et al. (2025): Improved Segmentation: More robust handling of typographic quotation marks during the sentence splitting process, leading to cleaner boundaries. Length Filtering: Implementation of a minimum fragment threshold of 10 characters to remove uninformative or noisy short segments. Strict Data Leakage Prevention: We explicitly verified to identify and remove any cross-split overlap. This ensures that the training, validation, and test sets are fully disjoint (Train $\cap$ Val $= \emptyset$, Train $\cap$ Test $= \emptyset$, Val $\cap$ Test $= \emptyset$). Verse-Level Alignment: Unlike corpora derived from noisy web crawling, this dataset uses verse-level synchronization, ensuring precise one-to-one correspondence between the source (Spanish) and target (Aymara) languages without the need for error-prone automatic alignment tools. Data Split Strategy The dataset is provided in a stratified 80/10/10 split: Train: 28,847 pairs Validation: 3,606 pairs Test: 3,606 pairs Technical Specifications Languages: Spanish (es-ES), Aymara (aym) Format: Plain text (.txt) / Tab-separated values (.tsv) Domain: Religious/Biblical (High-quality native translation) Citation and Attribution If you use this dataset, please cite the following work: TODO This work is an evolution of: Nina, M. and Vega-Oliveros, D. A. "Maximizing Model Adaptation for Low-Resource Languages: A Progressive Unfreezing Strategy for Spanish-Aymara Translation," 2025 12th International Conference on Soft Computing & Machine Intelligence (ISCMI), Rio de Janeiro, Brazil, 2025, pp. 259-263, doi: 10.1109/ISCMI67495.2025.11358562.



