Uyghur
收藏资源简介:
Uyghur AI Corpus 是一个精心策划的维吾尔语文本数据集,旨在为人工智能模型提供基础训练资源,使大语言模型能够理解、生成和翻译维吾尔语。数据集源自公开互联网,收集了多年来在各种公共平台、论坛和网站上分享的维吾尔族集体知识遗产,包括故事、散文、文章和历史记载等。原始网络数据经过仔细清洗、格式化并结构化处理,以满足高质量AI训练标准。数据集中尽可能地保留了元数据(如作者和来源),以尊重原始创作者。数据集包含以下字段:标题(title)、主要文本内容(text)、作者(author)、来源平台(source)、发布日期(date)以及译者(translator)。该数据集适用于文本生成、翻译和掩码填充等自然语言处理任务。最新更新(2026年2月)新增了维吾尔语文章和诗歌。数据集许可证为MIT。
Uyghur AI Corpus is a carefully curated Uyghur text dataset designed to provide foundational training resources for AI models, enabling large language models to understand, generate, and translate Uyghur. The dataset originates from public internet, collecting the collective knowledge heritage of the Uyghur people shared over the years on various public platforms, forums, and websites, including stories, essays, articles, and historical records. The original web data has been carefully cleaned, formatted, and structured to meet high-quality AI training standards. Metadata (such as author and source) are preserved as much as possible to respect the original creators. The dataset contains the following fields: title, text, author, source, date, and translator. It is suitable for natural language processing tasks such as text generation, translation, and masked language modeling. The latest update (February 2026) added Uyghur articles and poems. The dataset is licensed under MIT.
数据集概述
Uyghur Corpus (AI-Optimized) 是一个专为人工智能训练优化的维吾尔语语料库,旨在通过提供高质量文本数据,帮助大语言模型(LLMs)理解、生成和翻译维吾尔语。
基本信息
- 语言: 维吾尔语(ug)
- 许可证: MIT
- 任务类别: 文本生成(text-generation)、翻译(translation)、掩码填充(fill-mask)
- 标签: uyghur、corpus、nlp、uyghur-dataset、uyghur-ai
- 数据集主页: https://huggingface.co/datasets/Uyghur-Corpus/Uyghur
- 配置: 默认配置(default),包含一个训练集(train),数据文件格式为
*.jsonl
数据集内容与来源
该语料库是从开放的互联网上精心收集并整理的文本集合,代表了维吾尔民族在各类公共平台、论坛和网站上分享的集体知识遗产。来源类型包括:
- 故事(Stories)
- 散文(Essays)
- 文章(Articles)
- 历史记述(Historical accounts)
原始网页数据经过细致的清洗、格式化和结构化处理,以满足高质量AI训练的标准。在收集过程中,尽可能保留了原始创作者的 author(作者)和 source(来源)元数据,以尊重知识产权。
数据字段说明
| 字段名 | 含义 |
|---|---|
title |
作品标题 |
text |
训练使用的主要文本内容 |
author |
作品原作者 |
source |
作品来源平台 |
date |
发布日期 |
translator |
译者姓名 |
更新动态(2026年2月更新)
数据集近期进行了内容更新,新增了维吾尔语文章和诗歌。




