somosnlp-hackathon-2026/rosettia-chanka-data
收藏资源简介:
RosettIA Chanka Quechua — Judicial Parallel Data 是一个用于司法/行政管理领域的西班牙语与Chanka/Ayacucho Quechua(语言代码:quy)平行语料数据集。数据来源于秘鲁文化部2014年出版的手册《Manual para el empleo del Quechua Chanka en la administración de justicia》,该手册明确允许在注明出处的情况下复制。数据集包含多个Parquet文件,涵盖经过审查的西班牙语-Chanka平行对(包括斜杠替代拆分)、词汇表条目(如术语对),以及相关元数据。列包括reviewed_spanish和reviewed_chanka_quechua等。该数据集是RosettIA项目的一部分,专门用于翻译任务,支持Quechua语言资源开发。许可证为CC-BY-4.0,要求署名。注意:此发布仅包含可自由分发的数据,不包括其他受限制的第三方语料。
RosettIA Chanka Quechua — Judicial Parallel Data is a Spanish ↔ Chanka/Ayacucho Quechua (quy) parallel dataset for the judicial/administration of justice domain. It is derived from a Peruvian Ministry of Culture manual (2014) that explicitly permits reproduction with attribution. The dataset includes multiple Parquet files containing reviewed Spanish-Chanka parallel pairs (with slash-alternative splits), glossary entries (e.g., term pairs), and metadata. Columns include reviewed_spanish and reviewed_chanka_quechua. This dataset is part of the RosettIA project, aimed at translation tasks and Quechua language resource development. It is licensed under CC-BY-4.0, requiring attribution. Note: This release only includes the cleanly redistributable data, excluding other third-party corpora with restrictive licenses.




