遇见数据集

Shawi-Amazon: A Parallel Dataset for Spanish-Shawi Low-Resource Machine Translation

收藏
Zenodo2026-01-15 更新2026-05-26 收录
官方服务:

资源简介:

Low-resource languages, particularly those from the Amazonian region, remain largely underrepresented in current Natural Language Processing (NLP) research. In this work, we introduce the Shawi-Amazon Corpus, the first standardized parallel dataset for the Shawi (Chayahuita) language paired with Spanish. The corpus comprises approximately 9,210 aligned sentence pairs derived from the New Testament and Genesis. We detail a robust data engineering pipeline designed to address complex alignment challenges, specifically "many-to-one" verse mappings and textual variants between the Textus Receptus and Critical Text traditions. To ensure rigorous benchmarking, we implement a document-level splitting strategy, preventing data leakage between training and evaluation sets. This resource is released in standardized formats to facilitate future research in Neural Machine Translation (NMT) for the Cahuapanan language family, contributing to the digital preservation of indigenous heritage.

提供机构:
Zenodo
创建时间:
2026-01-15
二维码
社区交流群
二维码
科研交流群
商业服务