遇见数据集

Thermostatic/rosettia-chanka-data

收藏
Hugging Face2026-05-29 更新2026-05-31 收录
官方服务:

资源简介:

--- license: other language: - es - qu - quy task_categories: - translation tags: - quechua - chanka - spanish - parallel-corpus - rosettia - somosnlp-2026 pretty_name: RosettIA Chanka Quechua Parallel Data --- # RosettIA Chanka Quechua Parallel Data This dataset release supports the RosettIA - Quechua project for the SomosNLP 2026 hackathon. GitHub repository: https://github.com/Sekinal/rosettia-chanka ## What Is Included ### `clean_chanka/` Reviewed Spanish-Chanka Quechua artifacts extracted from *Manual para el empleo del Quechua Chanka en la administracion de justicia* (Ministerio de Cultura del Peru, 2014). Main training file: - `clean_chanka/manual_quechua_chanka_parallel_training_ready_augmented.parquet` - 1,055 reviewed Spanish-Chanka pairs. Other clean Chanka files: - `manual_quechua_chanka_parallel_reviewed.parquet` - 1,012 reviewed extraction rows with decisions/flags. - `manual_quechua_chanka_parallel_alternative_splits.parquet` - 81 rows created by splitting slash alternatives. - `manual_quechua_chanka_glossary_entries.parquet` - 423 glossary entries. - `manual_quechua_chanka_glossary_simple_terms.parquet` - 219 simple glossary term pairs. The source PDF is **not included** in this dataset repository. ### `broad_quechua/` Filtered broad Spanish-Quechua data derived from the public Hugging Face dataset `somosnlp-hackathon-2022/spanish-to-quechua`. Recommended broad SFT file: - `broad_quechua/somosnlp_spanish_to_quechua_high_quality_sft.parquet` - 82,873 strict no-flag filtered pairs. Ablation/debug files: - `somosnlp_spanish_to_quechua_broad_sft.parquet` - 111,699 broader backup filtered pairs. - `somosnlp_spanish_to_quechua_scored.parquet` - 123,839 unique normalized pairs with scores and filter flags. This broad tier is **not Chanka-verified**. Use it for initial broad Quechua-Spanish translation adaptation only, then fine-tune/evaluate on the clean Chanka tier. ### `americasnlp/` Filtered broad Quechua-Spanish data derived from AmericasNLP 2024 ST1 `quechua-spanish` resources. Recommended real-source file: - `americasnlp/americasnlp_quechua_spanish_high_quality_real_sft.parquet` - 87,004 strict high-quality real-source pairs. Evaluation/debug files: - `americasnlp_quechua_spanish_eval.parquet` - 996 dev/evaluation rows. - `americasnlp_quechua_spanish_scored.parquet` - 194,665 unique normalized rows with scores and filter flags. This tier includes `quy` data but is **not reviewed Chanka judicial-domain data**. Keep it separate from `clean_chanka/`. Synthetic/backtranslated AmericasNLP data is intentionally not included in this public release. ## Recommended Training Use 1. Broad SFT option A: `broad_quechua/somosnlp_spanish_to_quechua_high_quality_sft.parquet` 2. Broad SFT option B: `americasnlp/americasnlp_quechua_spanish_high_quality_real_sft.parquet` 3. Chanka/domain adaptation: `clean_chanka/manual_quechua_chanka_parallel_training_ready_augmented.parquet` Keep the two tiers separate in experiments. ## License And Provenance This is a mixed-provenance release, so the dataset card uses `license: other`. - The clean Chanka source document states reproduction is permitted when citing the source. The PDF itself is not uploaded here. - The broad Quechua files are derived from `somosnlp-hackathon-2022/spanish-to-quechua`; the visible dataset card did not expose a clear license during our review. Users should check upstream terms before redistribution or commercial use. See `metadata/source_manifest.csv` and the markdown files in `metadata/` for counts, caveats, and reproducibility notes. ## Reproduce Clone the GitHub repo and run: ```bash uv sync uv run python scripts/extract_manual_parallel_corpus.py uv run python scripts/review_manual_parallel_corpus.py uv run python scripts/split_manual_alternatives.py uv run python scripts/extract_manual_glossary.py uv run python scripts/preprocess_somosnlp_dirty_parallel.py ``` The raw source PDF is not included here; place it at the path documented in the GitHub repo before reproducing the manual extraction.

This dataset supports the RosettIA - Quechua project for the SomosNLP 2026 hackathon, providing parallel data between Spanish and Quechua (specifically the Chanka dialect). It includes three main parts: 1) clean_chanka/: reviewed Spanish-Chanka Quechua artifacts extracted from a 2014 Peruvian Ministry of Culture manual on judicial administration, containing 1,055 high-quality parallel pairs and glossary entries; 2) broad_quechua/: filtered broad Spanish-Quechua data derived from a public Hugging Face dataset, with 82,873 strict high-quality parallel pairs for initial translation adaptation; 3) americasnlp/: filtered broad Quechua-Spanish data from AmericasNLP 2024 resources, containing 87,004 high-quality real-source parallel pairs for evaluation and debugging. The dataset is designed for translation tasks, covering Spanish, Quechua Chanka (qu), and Quechua variants (quy), but does not include the original PDF file, with varying provenance and licensing across sections.

提供机构:
Thermostatic
搜集汇总
数据集介绍
Thermostatic/rosettia-chanka-data 数据集图片
构建方式
RosettIA Chanka Quechua — Judicial Parallel Data 数据集源自秘鲁文化部于2014年出版的《司法行政中Chanka Quechua语使用手册》,该手册明确允许在注明出处的前提下进行复制。研究团队从该手册中系统提取并审校了西班牙语与Chanka Quechua语(亦称Ayacucho Quechua语)之间的平行句对及词汇表,生成了涵盖1055条训练用平行句对、1012条审校记录、81条斜杠替代拆分句对以及642条词汇条目的结构化数据。所有数据均以Parquet格式存储,并附带详细的审校元数据与来源注释,确保数据来源的清晰与可追溯。
特点
该数据集聚焦于司法与行政领域的专业语言对,具有高度的领域专精性,并非通用领域的Chanka Quechua语料。其显著特点在于所有数据均可基于CC-BY-4.0许可协议干净地再分发,避免了第三方数据源的许可障碍。此外,数据集中包含了经过人工审校的平行句对与词汇条目,并提供了审校决策与标记信息,提升了数据的可靠性与实用价值。数据集的设计兼顾了学术研究与实际应用需求,为低资源语言的司法翻译任务提供了稀缺的高质量资源。
使用方法
用户可通过HuggingFace Datasets库直接加载clean_chanka目录下的Parquet文件,利用pandas或datasets API读取平行句对与词汇表,用于训练或评估西班牙语与Chanka Quechua语之间的机器翻译模型。建议将数据划分为训练集与验证集,结合其他合规的第三方语料库(如FLORES-200、OPUS等)进行数据增强。由于数据领域特定,用户在使用前应充分了解司法行政术语的背景,并依据CC-BY-4.0许可要求标注原始手册作者与RosettIA项目贡献者。训练脚本与数据处理流程可在项目的GitHub仓库中找到。
背景与挑战
背景概述
RosettIA Chanka Quechua — Judicial Parallel Data 数据集由 Estefanía Espinosa Fernández 与 Irving Ernesto Quezada Ramírez 于秘鲁文化部授权框架下创建,旨在构建司法领域西班牙语与昌卡克丘亚语(Ayacucho Quechua)之间的平行语料。该数据来源于 Wilfredo Ardito Vega 2014 年出版的《Manual para el empleo del Quechua Chanka en la administración de justicia》,经人工审校后提取出约千余条高质量平行句对及词典条目。作为 RosettIA 项目的核心开源组件,该数据集填补了低资源克丘亚语言在司法行政场景中机器翻译与自然语言处理研究的空白,为后续模型训练与评估提供了可复现的基准,对推动濒危语言数字化贡献显著。
当前挑战
该数据集面临的核心挑战包括:领域专业性导致的语料稀缺,司法文本术语复杂且对翻译准确性要求极高,而公开可用的克丘亚语平行语料极为有限;版权与分发限制严峻,多数第三方来源(如 JW300、FLORES-200、词典及网络爬取语料)因许可证冲突、非商业用途限制或未明确授权而无法纳入公开版本,仅能提供构建脚本;数据质量参差不齐,大规模合成数据需经蒸馏与强化学习过滤,而人工审校的少量数据又面临规模与领域覆盖不足的困境,模型泛化能力受限。
常用场景
经典使用场景
RosettIA Chanka Quechua数据集的核心用途在于构建与评估西班牙语与占卡克丘亚语之间的司法领域平行语料库,尤其聚焦于秘鲁司法行政中的双语翻译任务。该数据集整合了经人工审核的司法文本对与术语表,为低资源语言机器翻译提供了高质量的监督训练数据,支持模型在专业领域内实现精准的跨语言转换。研究者可借此探索在严格领域限定下,如何利用有限标注语料提升翻译系统的专业术语一致性、法律表述规范性以及上下文理解能力,为其他资源匮乏语言的专用翻译研究提供方法论参考。
解决学术问题
该数据集直面低资源语言机器翻译中的两大关键障碍:专业领域平行语料的匮乏与数据版权合规的复杂性。通过公开一份经版权许可、来源于秘鲁文化部官方手册的司法平行语料,它为研究者在受限数据条件下优化翻译模型性能铺平了道路。这一资源解决了验证数据增强、迁移学习、强化学习等策略在保护原住民语言翻译效果时缺乏可靠基准的问题,推动了学术界对克丘亚语等濒危语言在公共服务领域应用可行性的系统性评估,其影响延伸至语言技术公平性与文化保护交叉领域。
衍生相关工作
此数据集的构建催生了一系列衍生研究,包括基于NLLB模型进行知识蒸馏以生成合成平行语料的实验,以及利用Group-based Supervised Preference Optimization进行强化学习微调的探索。此外,研究者利用该语料对比了DoRA、rsLoRA与普通LoRA等参数高效微调方法在低资源翻译任务上的表现,并设计了多种数据混合策略以评估领域语料与通用语料联合训练的效果。这些工作共同推动了适用于司法场景的克丘亚语翻译系统性能边界,并为后续在安第斯地区其他土著语言上的模型适配提供了可复现的技术框架。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务