遇见数据集

nuosu-corpus

收藏
Hugging Face2026-08-02 更新2026-08-03 收录
官方服务:

资源简介:

诺苏语任务语料库(Nuosu Task-Oriented Corpus)是一个面向低资源语言诺苏语(标准凉山彝语)的学术语料库,由西安交通大学 Wuhe Axi 创建,版本 2026.08.02。该语料库按照模型训练目标组织数据,而非按来源收集,每条记录均保留来源、修订、文档、页面、转换和质量等元数据。数据集包含三个主要子集:ready_sft.jsonl(113,269条记录,用于监督微调)、ready_cpt.jsonl(4,238条记录,用于继续预训练)、normative_reference.jsonl(1,220条记录,作为Unicode和正字法参考)。另有 excluded.jsonl(55条记录,不含标准彝文文本,不用于训练)。数据来源包括彝汉社区词典、RFLR诗行对齐平行文本、Unicode和SIL字符/拼音/IPA映射、CLDR术语、UDHR、Apertium、Wiktionary、Tatoeba、OPUS、清洗后的维基百科和网页文本、可直接提取的PDF文本、OCR衍生的彝文文本,以及1,030个人工验证的OCR地面真值文本。每条记录带有上游归属和许可元数据,使用者需遵循各来源的条款。该语料库适用于文本生成和翻译任务,尤其适合低资源语言诺苏语的模型训练与评估。注意:语料库高度偏向短字典翻译示例,许多记录为机器衍生或社区贡献,基准重叠被有意保留,因此不能用于声称无污染的 NuosuBench 结果。

Nuosu Task-Oriented Corpus is an academic corpus targeted at the low-resource language Nuosu (Standard Liangshan Yi), created by Wuhe Axi from Xi'an Jiaotong University, with version 2026.08.02. This corpus structures data according to model training objectives rather than data collection sources. Each record retains metadata such as source, revision, document, page, conversion status and quality. The dataset contains three main subsets: ready_sft.jsonl (113,269 records for supervised fine-tuning), ready_cpt.jsonl (4,238 records for continued pre-training), and normative_reference.jsonl (1,220 records as Unicode and orthography reference). There is also an excluded.jsonl file containing 55 records that do not contain standard Yi script text and are not intended for training. Data sources include Yi-Chinese community dictionaries, RFLR verse-aligned parallel texts, Unicode and SIL character, pinyin and IPA mappings, CLDR terminology, UDHR, Apertium, Wiktionary, Tatoeba, OPUS, cleaned Wikipedia and web texts, directly extractable PDF texts, OCR-derived Yi script texts, and 1,030 manually verified OCR ground truth texts. Each record includes upstream attribution and license metadata, and users are required to comply with the terms of each original source. This corpus is suitable for text generation and translation tasks, and is especially useful for model training and evaluation of the low-resource language Nuosu. Note: The corpus is heavily biased towards short dictionary translation examples, many records are machine-generated or community-contributed, and benchmark overlaps are intentionally retained, so it cannot be used to claim pollution-free NuosuBench results.

创建时间:
2026-07-28
原始信息汇总

诺苏语任务语料库(Nuosu Task-Oriented Corpus)

数据集基本信息

  • 作者: Wuhe Axi,西安交通大学
  • 版本: 2026.08.02
  • 语言: 彝语(ii)、中文(zh)、英文(en)
  • 许可: 其他(other)
  • 规模: 100K < n < 1M
  • 任务类别: 文本生成(text-generation)、翻译(translation)
  • 标签: 诺苏语、四川彝语、凉山彝语、标准彝语、低资源语言

数据集结构与规模

数据条目 记录数 用途
ready_sft.jsonl 113,269 监督微调
ready_cpt.jsonl 4,238 持续预训练
normative_reference.jsonl 1,220 Unicode 和正字法参考
excluded.jsonl 55 无标准彝语文本的页面(不应被训练)

数据集中包含 1,030 条人工核验的 OCR 真值文本,已作为 CPT 文本全部并入 ready_cpt.jsonl。数据集排除了原始 NuosuBench 记录、yidir 与 OCR 图像,也不包含 manual_review.jsonl 等预留接口的训练数据。

数据组织方式

该语料库按模型训练目标组织,而非按采集来源组织。每条记录均保留来源、修订、文档、页面、转换及质量信息。task 字段描述语料用途,metadata 字段保留来源、许可、文档、页面、OCR 引擎及质量来源等溯源信息。

主要来源

语料库整合了以下来源:

  • 彝汉社区词典
  • RFLR 逐行对齐平行文本
  • Unicode 与 SIL 字符/拼音/IPA 映射
  • CLDR 术语、UDHR、Apertium、Wiktionary、Tatoeba、OPUS
  • 清洗后的维基百科与网页文本
  • 可直接提取的 PDF 文本
  • OCR 生成的彝语文本
  • 1,030 条人工核验的 OCR 真值文本

每条记录均保留其独立的上游出处与许可元数据。

引用信息

提供了 BibTeX 和 GB/T 7714 两种引用格式,并建议同时引用实际使用子集相关的上游来源。

局限性说明

  • 语料库严重偏向短词典翻译示例
  • 大量记录为机器生成或社区贡献
  • 基准测试重叠被有意保留,该数据集不可用于声称无污染的 NuosuBench 评测结果
二维码
社区交流群
二维码
科研交流群
商业服务